FORECAST-CAT S16 · Certified robustness, attack surface and eval-awareness

Benchmark forecasting from internals plus a 200-pair slice

Mixed

frozen 2026-07-18 · adjudicated 2026-07-19 · ev:1acdc14e2fa16124786a9cd0d44a03b2 · spec:39a99a5ab61d6

The frozen predictions

  1. H-forecast-mae internals plus a 200-pair slice forecast per-category RB2 accuracy within 0.06 MAE
    forecast_mae < 0.06 registered p 0.55 forecast_mae Value 0.08695840679817465 Subject FORECAST-CAT Threshold < 0.06 Trust Adjudicated Evidence ev:1acdc14e2fa16124786a9cd0d44a03b2 Refuted
  2. H-forecast-beats-size the forecast beats predict-from-size
    forecast_mae_minus_size_baseline < 0 forecast_mae_minus_size_baseline Value -0.10242024172821523 Subject FORECAST-CAT Threshold < 0 Trust Adjudicated Evidence ev:1acdc14e2fa16124786a9cd0d44a03b2 Confirmed
Adjudication figure for FORECAST-CAT
FORECAST-CAT against its frozen predictions: mixed.

Registered companions

Subjects and artifacts

  • signals: fleet
  • datasets: rb2-full , rb2-calibration-200
  • frozen spec: campaign-forecast-cat.json
  • figure files: light · dark
  • study id: study:campaign-forecast-cat@v1#39a99a5a