Supplemental field-trial laboratory
When the average hides a rare failure
A deterministic, prediction-first bench. Every control has units; every result has an inspectable table and static account.
Prediction
At five percent code prevalence, which candidate will prevalence-weighted mean risk select, and which candidate will worst-environment risk select? Explain why before revealing the bench.
Worst-environment risk ignores prevalence; a weighted mean can prefer the fragile subset below the crossover.
Practice only · this interaction never becomes mastery evidence.
Loading the verified local model…
Read the accessible static account
Rare room discounted
The frequency-first candidate's large code loss is multiplied by a small prevalence, so its weighted mean remains lower even though its worst-room loss is 0.84.
- code prevalence
- 5%
- fragile mean
- 0.1294
- mean selects
- Frequency-first
- robust mean
- 0.1579
- worst selects
- Robust capacity
Mean-risk crossover
At the crossover the candidates have the same prevalence-weighted mean risk. Their worst-environment risks remain 0.84 and 0.27.
- code prevalence
- 9.5238%
- mean risk
- 0.163238
- mean selects
- Tie
- worst selects
- Robust capacity
Rare room carries enough weight
Once code receives enough prevalence, its large frequency-first loss outweighs that candidate's small advantage in common environments.
- code prevalence
- 15%
- fragile mean
- 0.2042
- mean selects
- Robust capacity
- robust mean
- 0.1697
- worst selects
- Robust capacity
Concept anchors
The scenario connects scale, evidence, and model evaluation.
The lab is attached to concepts through explicit APPLIES relations. It does not alter their lesson containment or claim review state.
Continue the record
Move between intuition and evidence.
Ephemerent News tells the result as an intuitive argument. Ephemerent Research preserves the paper, review, files, and version history. Atlas keeps this conceptual scenario available for sustained manipulation.