PREDICTIONS.md:155
| R3 | `Lambda` moves **before** the labelled hack rate | **Resolved, and it lags.** On the real AISI series (`olmo3.1-32b-kl0.0-seed2`, 25,664 rollouts over 401 steps) the hack rate goes 0.002 to 0.989 with a fitted midpoint at step 106.0 and a 10-90 width of 23.9 steps, while `Lambda`'s own fitted midpoint is step 128.3. **`Lambda` lags by 0.93 transition widths.** The sign is robust and the magnitude is not: masking degenerate groups moves it to -3.6. `Lambda` itself is real on this series, 0.0371 [0.0128, 0.0789] against a permuted-step null median of 0.00062 at p = 0.001, and `eta_eff` is 6 sigma from zero, so what failed is the lead-time claim and not the instrument. **The registered baseline is unresolved and permanently so on this data**: the published rollout table carries no optimiser telemetry, so the gradient-norm comparator cannot be computed and would need the training log from the same run |

PREDICTIONS.md:154
| R4 | I5's variance derivative beats the gradient-norm peak on lead time | **Resolved in favour, on a real reward-hacking run.** `ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2`, 400 steps, 25,664 labelled rollouts, hack rate 0.016 to 0.984, transition midpoint step 106.0 and 10-90 width 23.9 steps at R2 0.996. **The gradient-norm peak arrives at step 133, which is 27 steps after the midpoint: a lead of -1.129 widths.** I5's derivative alarms at step 90, a lead of **+0.668 widths**. Both registered baselines are beaten and the second is beaten in a way the raw lead hides: the variance *level* alarms earlier still, at step 70 for +1.504 widths, and **fires on 60.7% of order-destroyed in-control surrogates from the same series**, as does the gradient norm under CUSUM at 69.3%. I5's derivative fires on 24.3%, or 5.3% using the trainer log's own `reward_std^2`. A lead that fires on two thirds of surrogates is not a detection, and separating those was the work |
