Certified robustness, attack surface and eval-awareness

Benchmark forecasting from internals plus a 200-pair slice

mixed registered

The reading

  1. internals plus a 200-pair slice forecast per-category RB2 accuracy within 0.06 MAE

    forecast_mae < 0.06 0.55 Benchmark forecasting from internals plus a 200-pair slice mixed registered < 0.06 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record › refuted

  2. the forecast beats predict-from-size

    forecast_mae_minus_size_baseline < 0 Benchmark forecasting from internals plus a 200-pair slice mixed registered confidence interval -0.11965441047419614 to -0.07469374333560999 < 0 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record › confirmed

FORECAST-CAT against its frozen predictions: mixed.
FORECAST-CAT against its frozen predictions: mixed.

Measurements

forecast_mae_minus_size_baseline_ci_low
Benchmark forecasting from internals plus a 200-pair slice mixed registered confidence interval -0.11965441047419614 to -0.07469374333560999 < 0 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record ›
forecast_mae_minus_size_baseline_ci_high
Benchmark forecasting from internals plus a 200-pair slice mixed registered confidence interval -0.11965441047419614 to -0.07469374333560999 < 0 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record ›
size_baseline_mae
Benchmark forecasting from internals plus a 200-pair slice mixed registered none registered The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. Nothing was registered against this value in advance, so it is a supporting measurement rather than a tested one. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record ›
slice_baseline_mae
Benchmark forecasting from internals plus a 200-pair slice mixed registered none registered The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. Nothing was registered against this value in advance, so it is a supporting measurement rather than a tested one. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record ›

Provenance

card
FORECAST-CAT
frozen_at
adjudicated_at
spec_hash
spec:39a99a5ab61d6f8e744a40b3804d21ab
study_id
study:campaign-forecast-cat@v1#39a99a5a
power_acceptance
P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent