The archive
A preregistered campaign, kept whole.
A set of hypotheses frozen with their kill criteria before the runs, adjudicated afterwards by machinery against a merged evidence store. It describes work done against the previous version of the library, and it is kept here in full, including the calls that went nowhere.
The ledger
53 hypotheses: 16 confirmed, 21 refuted, 16 inconclusive. In the order the campaign ran them.
- Error anticipation on RewardBench 2 internal features beat the exposed seen-pair margin at predicting held-out RB2 completion errors on Qwen3-8B Error anticipation on RewardBench 2 refuted registered confidence interval -0.004634508227426182 to 0.0023281954874443313 > 0.03 The frozen study "Error anticipation on RewardBench 2", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-8b and skywork-v2-llama31-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. delta-AUROC point call at its own 0.03 threshold under the label-shuffle refit null (null sd 0.0235, null pass rate at threshold 0.2013); the registered CI exclusion carries size control and deltas at or above the MDE confirm at 80 percent. Open the record › > 0.03 refuted registered
- Error anticipation on RewardBench 2 the same holds on Llama-8B (cross-family) Error anticipation on RewardBench 2 refuted registered confidence interval -0.005095063654071338 to -0.0018241777525145248 > 0.03 The frozen study "Error anticipation on RewardBench 2", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-8b and skywork-v2-llama31-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. delta-AUROC point call at its own 0.03 threshold under the label-shuffle refit null (null sd 0.0235, null pass rate at threshold 0.2013); the registered CI exclusion carries size control and deltas at or above the MDE confirm at 80 percent. Open the record › > 0.03 refuted registered
- Benchmark forecasting from internals plus a 200-pair slice internals plus a 200-pair slice forecast per-category RB2 accuracy within 0.06 MAE Benchmark forecasting from internals plus a 200-pair slice mixed registered < 0.06 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record › < 0.06 refuted registered
- Benchmark forecasting from internals plus a 200-pair slice the forecast beats predict-from-size Benchmark forecasting from internals plus a 200-pair slice mixed registered confidence interval -0.11965441047419614 to -0.07469374333560999 < 0 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record › < 0 confirmed registered
- Bias battery predicts RM-Bench Hard ranking internal style susceptibility predicts the fleet's RM-Bench Hard ranking not measured > 0.6 inconclusive registered
- Bias battery predicts RM-Bench Hard ranking the fleet-level correlation clears its permutation null not measured < 0.05 inconclusive registered
- Bias battery predicts RM-Bench Hard ranking the internal battery at least matches the behavioral degradation baseline not measured >= 0 inconclusive registered
- Tail index forecasts the best-of-k plateau on PPE the tail index predicts where best-of-k stops helping on the five PPE verifiable sets not measured > 0.3 inconclusive registered
- Tail index forecasts the best-of-k plateau on PPE the lower-quantile best-of-n aggregation beats the mean against human labels, reproducing the PPE finding not measured == 1 inconclusive registered
- Same-recipe scale ladder with frozen 8B extrapolation at least 2 of 3 capture-only internal quantities land in their pre-frozen 8B intervals Same-recipe scale ladder with frozen 8B extrapolation mixed registered >= 2 The frozen study "Same-recipe scale ladder with frozen 8B extrapolation", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b and skywork-v2-qwen3-1.7b and skywork-v2-qwen3-4b and skywork-v2-qwen3-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › >= 2 confirmed registered
- Same-recipe scale ladder with frozen 8B extrapolation the extrapolation beats a no-trend (equal-to-4B) baseline Same-recipe scale ladder with frozen 8B extrapolation mixed registered confidence interval -0.5311824800459731 to 2.4359016583798008 < 0 The frozen study "Same-recipe scale ladder with frozen 8B extrapolation", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b and skywork-v2-qwen3-1.7b and skywork-v2-qwen3-4b and skywork-v2-qwen3-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0 refuted registered
- Susceptibility chi predicts best-of-n feature drift on a real bank chi predicts realized best-of-n feature-drift ordering on a real bank Susceptibility chi predicts best-of-n feature drift on a real bank confirmed registered confidence interval 0.7500000000000002 to 1 > 0.3 The frozen study "Susceptibility chi predicts best-of-n feature drift on a real bank", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-8b and skywork-v2-llama31-8b, over ultrafeedback-bank. It does not establish that the same holds on another subject, another dataset or another run. split-bank estimator: chi and drift read disjoint prompt halves, so the null re-centers at zero, but the correlation's exchangeable unit is k = 6 features and the null still passes the frozen 0.3 threshold 0.28 of the time; the registered CI exclusion carries size control and only orderings near the synthetic calibration (0.958) clear the 0.91 MDE. Open the record › > 0.3 confirmed registered
- Susceptibility chi predicts best-of-n feature drift on a real bank chi beats a base-variance-only baseline Susceptibility chi predicts best-of-n feature drift on a real bank confirmed registered confidence interval 0.07142857142857151 to 0.5714285714285716 > 0 The frozen study "Susceptibility chi predicts best-of-n feature drift on a real bank", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-8b and skywork-v2-llama31-8b, over ultrafeedback-bank. It does not establish that the same holds on another subject, another dataset or another run. split-bank estimator: chi and drift read disjoint prompt halves, so the null re-centers at zero, but the correlation's exchangeable unit is k = 6 features and the null still passes the frozen 0.3 threshold 0.28 of the time; the registered CI exclusion carries size control and only orderings near the synthetic calibration (0.958) clear the 0.91 MDE. Open the record › > 0 confirmed registered
- The Goodhart hump located from a base-policy tail index an interior gold-reward hump exists in the proxy/gold best-of-n arm not measured == 1 inconclusive registered
- The Goodhart hump located from a base-policy tail index the tail-predicted hump location matches the observed one within 0.5 nats not measured < 0.5 inconclusive registered
- The two-fine-tunes gauge result, settled the raw cross-fine-tune reward cosine reproduces the v1 fixture near 0.005 not measured abs< 0.02 inconclusive registered
- The two-fine-tunes gauge result, settled the frame-fixed cosine is far larger than the raw cosine (co-rotation) not measured > 0.4 inconclusive registered
- The two-fine-tunes gauge result, settled the top residual-angle concept matches the largest behavioral delta (formality) not measured == 1 inconclusive registered
- Canonical cross-family angles predict pairwise disagreement the canonical angle between two RMs' reward directions predicts their held-out disagreement rate not measured > 0.6 inconclusive registered
- Canonical cross-family angles predict pairwise disagreement the correlation clears a model-clustered permutation null not measured < 0.05 inconclusive registered
- Weights-derived indices flag the exploitable hack family before any attack weights-derived indices flag the realized-most-exploitable hack family above a behavioral baseline not measured > 0 inconclusive registered
- Certified erasure removes the exploit without benchmark loss LEACE erasure of the flagged direction removes the exploit Certified erasure removes the exploit without benchmark loss mixed registered > 0.8 The frozen study "Certified erasure removes the exploit without benchmark loss", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on hackfore-flagged, over hack-probes and rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.8 confirmed registered
- Certified erasure removes the exploit without benchmark loss benchmark accuracy is retained through the erasure Certified erasure removes the exploit without benchmark loss mixed registered > -0.01 The frozen study "Certified erasure removes the exploit without benchmark loss", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on hackfore-flagged, over hack-probes and rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > -0.01 refuted registered
- Does the process reward model verify or style-read the PRM localizes ProcessBench error steps Does the process reward model verify or style-read refuted registered > 0.7 The frozen study "Does the process reward model verify or style-read", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on qwen-prm, over processbench-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.7 refuted registered
- Does the process reward model verify or style-read the correctness preference is anchored, not style-carried < 0.5 confirmed registered
- Verdict before critique on a real generative judge the verdict decoded at the pre-critique position matches the final verdict Verdict before critique on a real generative judge confirmed registered >= 0.9 The frozen study "Verdict before critique on a real generative judge", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-critic, over judge-pairs-1000. It does not establish that the same holds on another subject, another dataset or another run. Open the record › >= 0.9 confirmed registered
- Contested-direction loading predicts human rater disagreement contested-direction loading predicts per-item human rater disagreement not measured > 0.2 inconclusive registered
- Contested-direction loading predicts human rater disagreement single-model contested loading comes within 0.05 of the 5x-cost ensemble baseline not measured < 0.05 inconclusive registered
- Skepticism and receipt reliance predict who rewards fabricated receipts the two group contrasts separate graders as their planted values would Skepticism and receipt reliance predict who rewards fabricated receipts refuted registered > 0.5 The frozen study "Skepticism and receipt reliance predict who rewards fabricated receipts", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet-subset, over receipt-triples-1600. It does not establish that the same holds on another subject, another dataset or another run. fraction of positive grader contrasts at n = 5 graders: the frozen 0.5 sits at the null center, the point call passes under the null 0.51 of the time, and 80 percent confirmation needs per-grader positive-contrast probability 0.67, the s15-plant regime (calibration 0.993). Open the record › > 0.5 refuted registered
- Skepticism and receipt reliance predict who rewards fabricated receipts the omission-credulity gap is real on constructed triples Skepticism and receipt reliance predict who rewards fabricated receipts refuted registered > 0.2 The frozen study "Skepticism and receipt reliance predict who rewards fabricated receipts", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet-subset, over receipt-triples-1600. It does not establish that the same holds on another subject, another dataset or another run. fraction of positive grader contrasts at n = 5 graders: the frozen 0.5 sits at the null center, the point call passes under the null 0.51 of the time, and 80 percent confirmation needs per-grader positive-contrast probability 0.67, the s15-plant regime (calibration 0.993). Open the record › > 0.2 refuted registered
- Value-convergence excess across the fleet beats the capability-matched null value-convergence excess on real reward-model pairs beats the capability-matched random-utility null Value-convergence excess across the fleet beats the capability-matched null refuted registered confidence interval -0.12281662932807336 to -0.009192229234832329 > 0 The frozen study "Value-convergence excess across the fleet beats the capability-matched null", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-helpfulness-512. It does not establish that the same holds on another subject, another dataset or another run. the vce > 0 margin sits at the null center by construction: the point call passes under the null 0.48 of the time and its power at the threshold-alternative equals that rate; the alpha-exact p-value companion and the registered CI exclusion carry the card's evidence. Open the record › > 0 refuted registered
- Value-convergence excess across the fleet beats the capability-matched null the convergent-pair alignment exceeds the random-utility null at p < 0.05 Value-convergence excess across the fleet beats the capability-matched null refuted registered < 0.05 The frozen study "Value-convergence excess across the fleet beats the capability-matched null", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-helpfulness-512. It does not establish that the same holds on another subject, another dataset or another run. the vce > 0 margin sits at the null center by construction: the point call passes under the null 0.48 of the time and its power at the threshold-alternative equals that rate; the alpha-exact p-value companion and the registered CI exclusion carry the card's evidence. Open the record › < 0.05 confirmed registered
- The fleet decodes a property it does not price the fleet-median KUI gap for the star property exceeds its permuted null The fleet decodes a property it does not price refuted registered > 0.2 The frozen study "The fleet decodes a property it does not price", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.2 refuted registered
- The fleet decodes a property it does not price the gap clears its permutation null The fleet decodes a property it does not price refuted registered < 0.05 The frozen study "The fleet decodes a property it does not price", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.05 refuted registered
- The fleet decodes a property it does not price a priced control property shows no such gap The fleet decodes a property it does not price refuted registered < 0.1 The frozen study "The fleet decodes a property it does not price", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.1 refuted registered
- Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor dark reward grows with the criteria-to-effective-dimension ratio Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor mixed registered > 0.6 The frozen study "Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet and armorm, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.6 refuted registered
- Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor the fleet correlation clears its permutation null Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor mixed registered < 0.05 The frozen study "Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet and armorm, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.05 refuted registered
- Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor the ArmoRM 19-head coherence matrix respects the Welch floor Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor mixed registered >= 0 The frozen study "Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet and armorm, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › >= 0 confirmed registered
- The RM knows it is being tested, and steering along that direction inflates reward a benchmark-versus-organic direction is decodable The RM knows it is being tested, and steering along that direction inflates reward mixed registered > 0.6 The frozen study "The RM knows it is being tested, and steering along that direction inflates reward", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on evalaware-rms, over evalaware-paired and wildchat-organic and arena-organic. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.6 confirmed registered
- The RM knows it is being tested, and steering along that direction inflates reward steering along the direction inflates the score The RM knows it is being tested, and steering along that direction inflates reward mixed registered > 0.05 The frozen study "The RM knows it is being tested, and steering along that direction inflates reward", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on evalaware-rms, over evalaware-paired and wildchat-organic and arena-organic. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.05 refuted registered
- The RM knows it is being tested, and steering along that direction inflates reward the probe does not spuriously flag organic prompts (equivalence bound, an affirmative null) The RM knows it is being tested, and steering along that direction inflates reward mixed registered < 0.1 The frozen study "The RM knows it is being tested, and steering along that direction inflates reward", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on evalaware-rms, over evalaware-paired and wildchat-organic and arena-organic. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.1 refuted registered
- Intransitive preference mass no scalar reward can represent real k-wise preference data carries a computable intransitive fraction Intransitive preference mass no scalar reward can represent confirmed registered The frozen study "Intransitive preference mass no scalar reward can represent", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on nectar-tournaments and ultrafeedback-tournaments. The intransitive-mass reading is dead: zero cyclic triples in 186,378 triangles, 10,000 of 10,000 sampled tournaments acyclic, and the measured value sits on the encoding floor. The store still carries the adjudication as confirmed, which is why this cannot be left to the outcome mapping: it would render as supported. Open the record › > 0.03 confirmed registered
- Surface biases enter the reward direction before quality features surface biases reach half their final loading earlier than quality features Surface biases enter the reward direction before quality features mixed registered > 1 The frozen study "Surface biases enter the reward direction before quality features", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b, over ultrafeedback-train-20k. It does not establish that the same holds on another subject, another dataset or another run. entry-ratio point call at the frozen > 1.0, the null median: power equals the null pass rate 0.501, ratios at or above 2.60 confirm at 80 percent, and the s07 plant sits far above that; passes only through the explicit accept-underpowered path. Open the record › > 1 refuted registered
- Surface biases enter the reward direction before quality features the reward direction settles rather than wandering Surface biases enter the reward direction before quality features mixed registered == 1 The frozen study "Surface biases enter the reward direction before quality features", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b, over ultrafeedback-train-20k. It does not establish that the same holds on another subject, another dataset or another run. entry-ratio point call at the frozen > 1.0, the null median: power equals the null pass rate 0.501, ratios at or above 2.60 confirm at 80 percent, and the s07 plant sits far above that; passes only through the explicit accept-underpowered path. Open the record › == 1 confirmed registered
- Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk instrument scorecard AUCs on a real 0.6B trunk match the CPU-organism AUCs within 0.15 Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk refuted registered < 0.15 The frozen study "Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b, over caltransfer-organisms. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.15 refuted registered
- Attribution versus patching, adjudicated against planted ground truth patching's ranking matches the organism's marker-responsive components (the twin-activation-delta answer key) better than attribution's Attribution versus patching, adjudicated against planted ground truth refuted registered confidence interval -0.7735849056603774 to -0.4276729559748428 > 0 The frozen study "Attribution versus patching, adjudicated against planted ground truth", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b and skywork-v02, over adjavp-organism and rb2-patch-40. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0 refuted registered
- Do the indices carry information beyond size, accuracy, and lineage the surviving index clears its permutation null Do the indices carry information beyond size, accuracy, and lineage mixed registered < 0.05 The frozen study "Do the indices carry information beyond size, accuracy, and lineage", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.05 confirmed registered
- Do the indices carry information beyond size, accuracy, and lineage at least one load-bearing index retains a partial correlation above 0.3 after the size-and-accuracy control Do the indices carry information beyond size, accuracy, and lineage mixed registered > 0.3 The frozen study "Do the indices carry information beyond size, accuracy, and lineage", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.3 confirmed registered
- Do the indices carry information beyond size, accuracy, and lineage the surviving index also survives the within-lineage check Do the indices carry information beyond size, accuracy, and lineage mixed registered == 1 The frozen study "Do the indices carry information beyond size, accuracy, and lineage", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › == 1 refuted registered
- The campaign scores its own calibration the campaign's directional predictions, each assigned a pre-run probability, beat a coin The campaign scores its own calibration refuted registered < 0.25 The frozen study "The campaign scores its own calibration", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on the campaign's own record, rather than a model or a dataset. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.25 refuted registered
- The campaign scores its own calibration the nominal-0.8 prediction intervals cover the observed values near their nominal rate The campaign scores its own calibration refuted registered >= 0.6 The frozen study "The campaign scores its own calibration", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on the campaign's own record, rather than a model or a dataset. It does not establish that the same holds on another subject, another dataset or another run. Open the record › >= 0.6 confirmed registered
- A legibility frontier leaves a tacit residual a residual no short rubric captures A legibility frontier leaves a tacit residual confirmed registered > 0 The frozen study "A legibility frontier leaves a tacit residual", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-llama31-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0 confirmed registered
- The reward Hessian's flat subspace overlaps the flagged hackable direction the flat directions of the reward Hessian overlap the HACK-FORE-flagged direction more than a random direction does not measured > 0 inconclusive registered
The ledger and the headline set are different counts of different things. The figure below counts 17 headline calls. The two do not add together and neither is a score.
Every headline call against its own frozen threshold
- Error anticipation on RewardBench 2 Error anticipation on RewardBench 2 refuted registered confidence interval -0.004634508227426182 to 0.0023281954874443313 > 0.03 The frozen study "Error anticipation on RewardBench 2", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-8b and skywork-v2-llama31-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. delta-AUROC point call at its own 0.03 threshold under the label-shuffle refit null (null sd 0.0235, null pass rate at threshold 0.2013); the registered CI exclusion carries size control and deltas at or above the MDE confirm at 80 percent. Open the record › > 0.03 refuted
- Benchmark forecasting from internals plus a 200-pair slice Benchmark forecasting from internals plus a 200-pair slice mixed registered < 0.06 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record › < 0.06 refuted
- Same-recipe scale ladder with frozen 8B extrapolation Same-recipe scale ladder with frozen 8B extrapolation mixed registered >= 2 The frozen study "Same-recipe scale ladder with frozen 8B extrapolation", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b and skywork-v2-qwen3-1.7b and skywork-v2-qwen3-4b and skywork-v2-qwen3-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › >= 2 confirmed
- Susceptibility chi predicts best-of-n feature drift on a real bank Susceptibility chi predicts best-of-n feature drift on a real bank confirmed registered confidence interval 0.7500000000000002 to 1 > 0.3 The frozen study "Susceptibility chi predicts best-of-n feature drift on a real bank", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-8b and skywork-v2-llama31-8b, over ultrafeedback-bank. It does not establish that the same holds on another subject, another dataset or another run. split-bank estimator: chi and drift read disjoint prompt halves, so the null re-centers at zero, but the correlation's exchangeable unit is k = 6 features and the null still passes the frozen 0.3 threshold 0.28 of the time; the registered CI exclusion carries size control and only orderings near the synthetic calibration (0.958) clear the 0.91 MDE. Open the record › > 0.3 confirmed
- Certified erasure removes the exploit without benchmark loss Certified erasure removes the exploit without benchmark loss mixed registered > 0.8 The frozen study "Certified erasure removes the exploit without benchmark loss", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on hackfore-flagged, over hack-probes and rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.8 confirmed
- Does the process reward model verify or style-read Does the process reward model verify or style-read refuted registered > 0.7 The frozen study "Does the process reward model verify or style-read", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on qwen-prm, over processbench-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.7 refuted
- Verdict before critique on a real generative judge Verdict before critique on a real generative judge confirmed registered >= 0.9 The frozen study "Verdict before critique on a real generative judge", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-critic, over judge-pairs-1000. It does not establish that the same holds on another subject, another dataset or another run. Open the record › >= 0.9 confirmed
- Skepticism and receipt reliance predict who rewards fabricated receipts Skepticism and receipt reliance predict who rewards fabricated receipts refuted registered > 0.5 The frozen study "Skepticism and receipt reliance predict who rewards fabricated receipts", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet-subset, over receipt-triples-1600. It does not establish that the same holds on another subject, another dataset or another run. fraction of positive grader contrasts at n = 5 graders: the frozen 0.5 sits at the null center, the point call passes under the null 0.51 of the time, and 80 percent confirmation needs per-grader positive-contrast probability 0.67, the s15-plant regime (calibration 0.993). Open the record › > 0.5 refuted
- Value-convergence excess across the fleet beats the capability-matched null Value-convergence excess across the fleet beats the capability-matched null refuted registered confidence interval -0.12281662932807336 to -0.009192229234832329 > 0 The frozen study "Value-convergence excess across the fleet beats the capability-matched null", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-helpfulness-512. It does not establish that the same holds on another subject, another dataset or another run. the vce > 0 margin sits at the null center by construction: the point call passes under the null 0.48 of the time and its power at the threshold-alternative equals that rate; the alpha-exact p-value companion and the registered CI exclusion carry the card's evidence. Open the record › > 0 refuted
- The fleet decodes a property it does not price The fleet decodes a property it does not price refuted registered > 0.2 The frozen study "The fleet decodes a property it does not price", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.2 refuted
- Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor mixed registered > 0.6 The frozen study "Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet and armorm, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.6 refuted
- The RM knows it is being tested, and steering along that direction inflates reward The RM knows it is being tested, and steering along that direction inflates reward mixed registered > 0.6 The frozen study "The RM knows it is being tested, and steering along that direction inflates reward", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on evalaware-rms, over evalaware-paired and wildchat-organic and arena-organic. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.6 confirmed
- Intransitive preference mass no scalar reward can represent Intransitive preference mass no scalar reward can represent confirmed registered The frozen study "Intransitive preference mass no scalar reward can represent", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on nectar-tournaments and ultrafeedback-tournaments. The intransitive-mass reading is dead: zero cyclic triples in 186,378 triangles, 10,000 of 10,000 sampled tournaments acyclic, and the measured value sits on the encoding floor. The store still carries the adjudication as confirmed, which is why this cannot be left to the outcome mapping: it would render as supported. Open the record › > 0.03 confirmed
- Surface biases enter the reward direction before quality features Surface biases enter the reward direction before quality features mixed registered > 1 The frozen study "Surface biases enter the reward direction before quality features", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b, over ultrafeedback-train-20k. It does not establish that the same holds on another subject, another dataset or another run. entry-ratio point call at the frozen > 1.0, the null median: power equals the null pass rate 0.501, ratios at or above 2.60 confirm at 80 percent, and the s07 plant sits far above that; passes only through the explicit accept-underpowered path. Open the record › > 1 refuted
- Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk refuted registered < 0.15 The frozen study "Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b, over caltransfer-organisms. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.15 refuted
- Attribution versus patching, adjudicated against planted ground truth Attribution versus patching, adjudicated against planted ground truth refuted registered confidence interval -0.7735849056603774 to -0.4276729559748428 > 0 The frozen study "Attribution versus patching, adjudicated against planted ground truth", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b and skywork-v02, over adjavp-organism and rb2-patch-40. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0 refuted
- Do the indices carry information beyond size, accuracy, and lineage Do the indices carry information beyond size, accuracy, and lineage mixed registered < 0.05 The frozen study "Do the indices carry information beyond size, accuracy, and lineage", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.05 confirmed
What this got wrong
The campaign reported a rank correlation of -0.256. The value measured from the fixture is -0.171, and it supersedes it. The old number is still live in the shipped README. It is written out here because the correction is the content, and this is the only route on the site where that string is allowed to appear.
- -0.256
- reward-lens-clean/docs/BUILD_NOTES.md:44
- -0.171
- reward-lens-clean/docs/BUILD_NOTES.md:50
The campaign scored its own calibration
It also asked whether its own directional predictions had been any good, and adjudicated that question by the same machinery it used on everything else. The answer, and the numbers behind it, are in the third act.
The cards
All 27 are here, grouped by verdict. The inconclusive ones are kept because removing them would be selection, and this is the page where selection would be least defensible.
confirmed · 4
refuted · 8
- Error anticipation on RewardBench 2
- Does the process reward model verify or style-read
- Skepticism and receipt reliance predict who rewards fabricated receipts
- Value-convergence excess across the fleet beats the capability-matched null
- The fleet decodes a property it does not price
- Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk
- Attribution versus patching, adjudicated against planted ground truth
- The campaign scores its own calibration
mixed · 7
- Benchmark forecasting from internals plus a 200-pair slice
- Same-recipe scale ladder with frozen 8B extrapolation
- Certified erasure removes the exploit without benchmark loss
- Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor
- The RM knows it is being tested, and steering along that direction inflates reward
- Surface biases enter the reward direction before quality features
- Do the indices carry information beyond size, accuracy, and lineage
inconclusive · 8
- Bias battery predicts RM-Bench Hard ranking
- Tail index forecasts the best-of-k plateau on PPE
- The Goodhart hump located from a base-policy tail index
- The two-fine-tunes gauge result, settled
- Canonical cross-family angles predict pairwise disagreement
- Weights-derived indices flag the exploitable hack family before any attack
- Contested-direction loading predicts human rater disagreement
- The reward Hessian's flat subspace overlaps the flagged hackable direction