Certified robustness, attack surface and eval-awareness
Benchmark forecasting from internals plus a 200-pair slice
mixed registered
The reading
-
internals plus a 200-pair slice forecast per-category RB2 accuracy within 0.06 MAE
forecast_mae < 0.06 0.55 Benchmark forecasting from internals plus a 200-pair slice mixed registered < 0.06 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record › refuted
-
the forecast beats predict-from-size
forecast_mae_minus_size_baseline < 0 Benchmark forecasting from internals plus a 200-pair slice mixed registered confidence interval -0.11965441047419614 to -0.07469374333560999 < 0 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record › confirmed
Measurements
- forecast_mae_minus_size_baseline_ci_low
- Benchmark forecasting from internals plus a 200-pair slice mixed registered confidence interval -0.11965441047419614 to -0.07469374333560999 < 0 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record ›
- forecast_mae_minus_size_baseline_ci_high
- Benchmark forecasting from internals plus a 200-pair slice mixed registered confidence interval -0.11965441047419614 to -0.07469374333560999 < 0 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record ›
- size_baseline_mae
- Benchmark forecasting from internals plus a 200-pair slice mixed registered none registered The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. Nothing was registered against this value in advance, so it is a supporting measurement rather than a tested one. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record ›
- slice_baseline_mae
- Benchmark forecasting from internals plus a 200-pair slice mixed registered none registered The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. Nothing was registered against this value in advance, so it is a supporting measurement rather than a tested one. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record ›
Provenance
- card
- FORECAST-CAT
- frozen_at
- adjudicated_at
- spec_hash
- spec:39a99a5ab61d6f8e744a40b3804d21ab
- study_id
- study:campaign-forecast-cat@v1#39a99a5a
- Evidence records
- 1acdc14e2fa16124786a9cd0d44a03b2
- power_acceptance
- P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent
- spec_file
- campaign-forecast-cat.json