The ledger

53 hypotheses: 16 confirmed, 21 refuted, 16 inconclusive. In the order the campaign ran them.

  1. Error anticipation on RewardBench 2 internal features beat the exposed seen-pair margin at predicting held-out RB2 completion errors on Qwen3-8B Error anticipation on RewardBench 2 refuted registered confidence interval -0.004634508227426182 to 0.0023281954874443313 > 0.03 The frozen study "Error anticipation on RewardBench 2", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-8b and skywork-v2-llama31-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. delta-AUROC point call at its own 0.03 threshold under the label-shuffle refit null (null sd 0.0235, null pass rate at threshold 0.2013); the registered CI exclusion carries size control and deltas at or above the MDE confirm at 80 percent. Open the record › > 0.03 refuted registered
  2. Error anticipation on RewardBench 2 the same holds on Llama-8B (cross-family) Error anticipation on RewardBench 2 refuted registered confidence interval -0.005095063654071338 to -0.0018241777525145248 > 0.03 The frozen study "Error anticipation on RewardBench 2", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-8b and skywork-v2-llama31-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. delta-AUROC point call at its own 0.03 threshold under the label-shuffle refit null (null sd 0.0235, null pass rate at threshold 0.2013); the registered CI exclusion carries size control and deltas at or above the MDE confirm at 80 percent. Open the record › > 0.03 refuted registered
  3. Benchmark forecasting from internals plus a 200-pair slice internals plus a 200-pair slice forecast per-category RB2 accuracy within 0.06 MAE Benchmark forecasting from internals plus a 200-pair slice mixed registered < 0.06 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record › < 0.06 refuted registered
  4. Benchmark forecasting from internals plus a 200-pair slice the forecast beats predict-from-size Benchmark forecasting from internals plus a 200-pair slice mixed registered confidence interval -0.11965441047419614 to -0.07469374333560999 < 0 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record › < 0 confirmed registered
  5. Bias battery predicts RM-Bench Hard ranking internal style susceptibility predicts the fleet's RM-Bench Hard ranking not measured > 0.6 inconclusive registered
  6. Bias battery predicts RM-Bench Hard ranking the fleet-level correlation clears its permutation null not measured < 0.05 inconclusive registered
  7. Bias battery predicts RM-Bench Hard ranking the internal battery at least matches the behavioral degradation baseline not measured >= 0 inconclusive registered
  8. Tail index forecasts the best-of-k plateau on PPE the tail index predicts where best-of-k stops helping on the five PPE verifiable sets not measured > 0.3 inconclusive registered
  9. Tail index forecasts the best-of-k plateau on PPE the lower-quantile best-of-n aggregation beats the mean against human labels, reproducing the PPE finding not measured == 1 inconclusive registered
  10. Same-recipe scale ladder with frozen 8B extrapolation at least 2 of 3 capture-only internal quantities land in their pre-frozen 8B intervals Same-recipe scale ladder with frozen 8B extrapolation mixed registered >= 2 The frozen study "Same-recipe scale ladder with frozen 8B extrapolation", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b and skywork-v2-qwen3-1.7b and skywork-v2-qwen3-4b and skywork-v2-qwen3-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › >= 2 confirmed registered
  11. Same-recipe scale ladder with frozen 8B extrapolation the extrapolation beats a no-trend (equal-to-4B) baseline Same-recipe scale ladder with frozen 8B extrapolation mixed registered confidence interval -0.5311824800459731 to 2.4359016583798008 < 0 The frozen study "Same-recipe scale ladder with frozen 8B extrapolation", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b and skywork-v2-qwen3-1.7b and skywork-v2-qwen3-4b and skywork-v2-qwen3-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0 refuted registered
  12. Susceptibility chi predicts best-of-n feature drift on a real bank chi predicts realized best-of-n feature-drift ordering on a real bank Susceptibility chi predicts best-of-n feature drift on a real bank confirmed registered confidence interval 0.7500000000000002 to 1 > 0.3 The frozen study "Susceptibility chi predicts best-of-n feature drift on a real bank", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-8b and skywork-v2-llama31-8b, over ultrafeedback-bank. It does not establish that the same holds on another subject, another dataset or another run. split-bank estimator: chi and drift read disjoint prompt halves, so the null re-centers at zero, but the correlation's exchangeable unit is k = 6 features and the null still passes the frozen 0.3 threshold 0.28 of the time; the registered CI exclusion carries size control and only orderings near the synthetic calibration (0.958) clear the 0.91 MDE. Open the record › > 0.3 confirmed registered
  13. Susceptibility chi predicts best-of-n feature drift on a real bank chi beats a base-variance-only baseline Susceptibility chi predicts best-of-n feature drift on a real bank confirmed registered confidence interval 0.07142857142857151 to 0.5714285714285716 > 0 The frozen study "Susceptibility chi predicts best-of-n feature drift on a real bank", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-8b and skywork-v2-llama31-8b, over ultrafeedback-bank. It does not establish that the same holds on another subject, another dataset or another run. split-bank estimator: chi and drift read disjoint prompt halves, so the null re-centers at zero, but the correlation's exchangeable unit is k = 6 features and the null still passes the frozen 0.3 threshold 0.28 of the time; the registered CI exclusion carries size control and only orderings near the synthetic calibration (0.958) clear the 0.91 MDE. Open the record › > 0 confirmed registered
  14. The Goodhart hump located from a base-policy tail index an interior gold-reward hump exists in the proxy/gold best-of-n arm not measured == 1 inconclusive registered
  15. The Goodhart hump located from a base-policy tail index the tail-predicted hump location matches the observed one within 0.5 nats not measured < 0.5 inconclusive registered
  16. The two-fine-tunes gauge result, settled the raw cross-fine-tune reward cosine reproduces the v1 fixture near 0.005 not measured abs< 0.02 inconclusive registered
  17. The two-fine-tunes gauge result, settled the frame-fixed cosine is far larger than the raw cosine (co-rotation) not measured > 0.4 inconclusive registered
  18. The two-fine-tunes gauge result, settled the top residual-angle concept matches the largest behavioral delta (formality) not measured == 1 inconclusive registered
  19. Canonical cross-family angles predict pairwise disagreement the canonical angle between two RMs' reward directions predicts their held-out disagreement rate not measured > 0.6 inconclusive registered
  20. Canonical cross-family angles predict pairwise disagreement the correlation clears a model-clustered permutation null not measured < 0.05 inconclusive registered
  21. Weights-derived indices flag the exploitable hack family before any attack weights-derived indices flag the realized-most-exploitable hack family above a behavioral baseline not measured > 0 inconclusive registered
  22. Certified erasure removes the exploit without benchmark loss LEACE erasure of the flagged direction removes the exploit Certified erasure removes the exploit without benchmark loss mixed registered > 0.8 The frozen study "Certified erasure removes the exploit without benchmark loss", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on hackfore-flagged, over hack-probes and rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.8 confirmed registered
  23. Certified erasure removes the exploit without benchmark loss benchmark accuracy is retained through the erasure Certified erasure removes the exploit without benchmark loss mixed registered > -0.01 The frozen study "Certified erasure removes the exploit without benchmark loss", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on hackfore-flagged, over hack-probes and rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > -0.01 refuted registered
  24. Does the process reward model verify or style-read the PRM localizes ProcessBench error steps Does the process reward model verify or style-read refuted registered > 0.7 The frozen study "Does the process reward model verify or style-read", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on qwen-prm, over processbench-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.7 refuted registered
  25. Does the process reward model verify or style-read the correctness preference is anchored, not style-carried Does the process reward model verify or style-read refuted registered < 0.5 The frozen study "Does the process reward model verify or style-read", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on qwen-prm, over processbench-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.5 confirmed registered
  26. Verdict before critique on a real generative judge the verdict decoded at the pre-critique position matches the final verdict Verdict before critique on a real generative judge confirmed registered >= 0.9 The frozen study "Verdict before critique on a real generative judge", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-critic, over judge-pairs-1000. It does not establish that the same holds on another subject, another dataset or another run. Open the record › >= 0.9 confirmed registered
  27. Contested-direction loading predicts human rater disagreement contested-direction loading predicts per-item human rater disagreement not measured > 0.2 inconclusive registered
  28. Contested-direction loading predicts human rater disagreement single-model contested loading comes within 0.05 of the 5x-cost ensemble baseline not measured < 0.05 inconclusive registered
  29. Skepticism and receipt reliance predict who rewards fabricated receipts the two group contrasts separate graders as their planted values would Skepticism and receipt reliance predict who rewards fabricated receipts refuted registered > 0.5 The frozen study "Skepticism and receipt reliance predict who rewards fabricated receipts", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet-subset, over receipt-triples-1600. It does not establish that the same holds on another subject, another dataset or another run. fraction of positive grader contrasts at n = 5 graders: the frozen 0.5 sits at the null center, the point call passes under the null 0.51 of the time, and 80 percent confirmation needs per-grader positive-contrast probability 0.67, the s15-plant regime (calibration 0.993). Open the record › > 0.5 refuted registered
  30. Skepticism and receipt reliance predict who rewards fabricated receipts the omission-credulity gap is real on constructed triples Skepticism and receipt reliance predict who rewards fabricated receipts refuted registered > 0.2 The frozen study "Skepticism and receipt reliance predict who rewards fabricated receipts", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet-subset, over receipt-triples-1600. It does not establish that the same holds on another subject, another dataset or another run. fraction of positive grader contrasts at n = 5 graders: the frozen 0.5 sits at the null center, the point call passes under the null 0.51 of the time, and 80 percent confirmation needs per-grader positive-contrast probability 0.67, the s15-plant regime (calibration 0.993). Open the record › > 0.2 refuted registered
  31. Value-convergence excess across the fleet beats the capability-matched null value-convergence excess on real reward-model pairs beats the capability-matched random-utility null Value-convergence excess across the fleet beats the capability-matched null refuted registered confidence interval -0.12281662932807336 to -0.009192229234832329 > 0 The frozen study "Value-convergence excess across the fleet beats the capability-matched null", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-helpfulness-512. It does not establish that the same holds on another subject, another dataset or another run. the vce > 0 margin sits at the null center by construction: the point call passes under the null 0.48 of the time and its power at the threshold-alternative equals that rate; the alpha-exact p-value companion and the registered CI exclusion carry the card's evidence. Open the record › > 0 refuted registered
  32. Value-convergence excess across the fleet beats the capability-matched null the convergent-pair alignment exceeds the random-utility null at p < 0.05 Value-convergence excess across the fleet beats the capability-matched null refuted registered < 0.05 The frozen study "Value-convergence excess across the fleet beats the capability-matched null", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-helpfulness-512. It does not establish that the same holds on another subject, another dataset or another run. the vce > 0 margin sits at the null center by construction: the point call passes under the null 0.48 of the time and its power at the threshold-alternative equals that rate; the alpha-exact p-value companion and the registered CI exclusion carry the card's evidence. Open the record › < 0.05 confirmed registered
  33. The fleet decodes a property it does not price the fleet-median KUI gap for the star property exceeds its permuted null The fleet decodes a property it does not price refuted registered > 0.2 The frozen study "The fleet decodes a property it does not price", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.2 refuted registered
  34. The fleet decodes a property it does not price the gap clears its permutation null The fleet decodes a property it does not price refuted registered < 0.05 The frozen study "The fleet decodes a property it does not price", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.05 refuted registered
  35. The fleet decodes a property it does not price a priced control property shows no such gap The fleet decodes a property it does not price refuted registered < 0.1 The frozen study "The fleet decodes a property it does not price", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.1 refuted registered
  36. Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor dark reward grows with the criteria-to-effective-dimension ratio Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor mixed registered > 0.6 The frozen study "Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet and armorm, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.6 refuted registered
  37. Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor the fleet correlation clears its permutation null Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor mixed registered < 0.05 The frozen study "Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet and armorm, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.05 refuted registered
  38. Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor the ArmoRM 19-head coherence matrix respects the Welch floor Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor mixed registered >= 0 The frozen study "Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet and armorm, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › >= 0 confirmed registered
  39. The RM knows it is being tested, and steering along that direction inflates reward a benchmark-versus-organic direction is decodable The RM knows it is being tested, and steering along that direction inflates reward mixed registered > 0.6 The frozen study "The RM knows it is being tested, and steering along that direction inflates reward", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on evalaware-rms, over evalaware-paired and wildchat-organic and arena-organic. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.6 confirmed registered
  40. The RM knows it is being tested, and steering along that direction inflates reward steering along the direction inflates the score The RM knows it is being tested, and steering along that direction inflates reward mixed registered > 0.05 The frozen study "The RM knows it is being tested, and steering along that direction inflates reward", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on evalaware-rms, over evalaware-paired and wildchat-organic and arena-organic. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.05 refuted registered
  41. The RM knows it is being tested, and steering along that direction inflates reward the probe does not spuriously flag organic prompts (equivalence bound, an affirmative null) The RM knows it is being tested, and steering along that direction inflates reward mixed registered < 0.1 The frozen study "The RM knows it is being tested, and steering along that direction inflates reward", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on evalaware-rms, over evalaware-paired and wildchat-organic and arena-organic. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.1 refuted registered
  42. Intransitive preference mass no scalar reward can represent real k-wise preference data carries a computable intransitive fraction Intransitive preference mass no scalar reward can represent confirmed registered The frozen study "Intransitive preference mass no scalar reward can represent", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on nectar-tournaments and ultrafeedback-tournaments. The intransitive-mass reading is dead: zero cyclic triples in 186,378 triangles, 10,000 of 10,000 sampled tournaments acyclic, and the measured value sits on the encoding floor. The store still carries the adjudication as confirmed, which is why this cannot be left to the outcome mapping: it would render as supported. Open the record › > 0.03 confirmed registered
  43. Surface biases enter the reward direction before quality features surface biases reach half their final loading earlier than quality features Surface biases enter the reward direction before quality features mixed registered > 1 The frozen study "Surface biases enter the reward direction before quality features", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b, over ultrafeedback-train-20k. It does not establish that the same holds on another subject, another dataset or another run. entry-ratio point call at the frozen > 1.0, the null median: power equals the null pass rate 0.501, ratios at or above 2.60 confirm at 80 percent, and the s07 plant sits far above that; passes only through the explicit accept-underpowered path. Open the record › > 1 refuted registered
  44. Surface biases enter the reward direction before quality features the reward direction settles rather than wandering Surface biases enter the reward direction before quality features mixed registered == 1 The frozen study "Surface biases enter the reward direction before quality features", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b, over ultrafeedback-train-20k. It does not establish that the same holds on another subject, another dataset or another run. entry-ratio point call at the frozen > 1.0, the null median: power equals the null pass rate 0.501, ratios at or above 2.60 confirm at 80 percent, and the s07 plant sits far above that; passes only through the explicit accept-underpowered path. Open the record › == 1 confirmed registered
  45. Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk instrument scorecard AUCs on a real 0.6B trunk match the CPU-organism AUCs within 0.15 Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk refuted registered < 0.15 The frozen study "Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b, over caltransfer-organisms. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.15 refuted registered
  46. Attribution versus patching, adjudicated against planted ground truth patching's ranking matches the organism's marker-responsive components (the twin-activation-delta answer key) better than attribution's Attribution versus patching, adjudicated against planted ground truth refuted registered confidence interval -0.7735849056603774 to -0.4276729559748428 > 0 The frozen study "Attribution versus patching, adjudicated against planted ground truth", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b and skywork-v02, over adjavp-organism and rb2-patch-40. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0 refuted registered
  47. Do the indices carry information beyond size, accuracy, and lineage the surviving index clears its permutation null Do the indices carry information beyond size, accuracy, and lineage mixed registered < 0.05 The frozen study "Do the indices carry information beyond size, accuracy, and lineage", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.05 confirmed registered
  48. Do the indices carry information beyond size, accuracy, and lineage at least one load-bearing index retains a partial correlation above 0.3 after the size-and-accuracy control Do the indices carry information beyond size, accuracy, and lineage mixed registered > 0.3 The frozen study "Do the indices carry information beyond size, accuracy, and lineage", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.3 confirmed registered
  49. Do the indices carry information beyond size, accuracy, and lineage the surviving index also survives the within-lineage check Do the indices carry information beyond size, accuracy, and lineage mixed registered == 1 The frozen study "Do the indices carry information beyond size, accuracy, and lineage", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › == 1 refuted registered
  50. The campaign scores its own calibration the campaign's directional predictions, each assigned a pre-run probability, beat a coin The campaign scores its own calibration refuted registered < 0.25 The frozen study "The campaign scores its own calibration", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on the campaign's own record, rather than a model or a dataset. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.25 refuted registered
  51. The campaign scores its own calibration the nominal-0.8 prediction intervals cover the observed values near their nominal rate The campaign scores its own calibration refuted registered >= 0.6 The frozen study "The campaign scores its own calibration", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on the campaign's own record, rather than a model or a dataset. It does not establish that the same holds on another subject, another dataset or another run. Open the record › >= 0.6 confirmed registered
  52. A legibility frontier leaves a tacit residual a residual no short rubric captures A legibility frontier leaves a tacit residual confirmed registered > 0 The frozen study "A legibility frontier leaves a tacit residual", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-llama31-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0 confirmed registered
  53. The reward Hessian's flat subspace overlaps the flagged hackable direction the flat directions of the reward Hessian overlap the HACK-FORE-flagged direction more than a random direction does not measured > 0 inconclusive registered

The ledger and the headline set are different counts of different things. The figure below counts 17 headline calls. The two do not add together and neither is a score.

Every headline call against its own frozen threshold

-4 -2 0 2 4 Error anticipation on RewardBench 2 refuted Benchmark forecasting from internals plus a 200-pair slice refuted Same-recipe scale ladder with frozen 8B extrapolation confirmed Susceptibility chi predicts best-of-n feature drift on a real bank confirmed Certified erasure removes the exploit without benchmark loss confirmed Does the process reward model verify or style-read refuted Verdict before critique on a real generative judge confirmed Skepticism and receipt reliance predict who rewards fabricated receipts refuted Value-convergence excess across the fleet beats the capability-matched null refuted The fleet decodes a property it does not price refuted Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor refuted The RM knows it is being tested, and steering along that direction inflates reward confirmed Intransitive preference mass no scalar reward can represent confirmed Surface biases enter the reward direction before quality features refuted Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk refuted Attribution versus patching, adjudicated against planted ground truth refuted Do the indices carry information beyond size, accuracy, and lineage confirmed
  1. Error anticipation on RewardBench 2 Error anticipation on RewardBench 2 refuted registered confidence interval -0.004634508227426182 to 0.0023281954874443313 > 0.03 The frozen study "Error anticipation on RewardBench 2", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-8b and skywork-v2-llama31-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. delta-AUROC point call at its own 0.03 threshold under the label-shuffle refit null (null sd 0.0235, null pass rate at threshold 0.2013); the registered CI exclusion carries size control and deltas at or above the MDE confirm at 80 percent. Open the record › > 0.03 refuted
  2. Benchmark forecasting from internals plus a 200-pair slice Benchmark forecasting from internals plus a 200-pair slice mixed registered < 0.06 The frozen study "Benchmark forecasting from internals plus a 200-pair slice", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-calibration-200. It does not establish that the same holds on another subject, another dataset or another run. P(observed MAE beats the no-skill null's 5 percent boundary 0.048 | true MAE at the frozen 0.06), 6 categories; a no-skill forecast passes the frozen bound itself 0.100 of the time, and true MAEs at or below 0.038 detect at 80 percent. Open the record › < 0.06 refuted
  3. Same-recipe scale ladder with frozen 8B extrapolation Same-recipe scale ladder with frozen 8B extrapolation mixed registered >= 2 The frozen study "Same-recipe scale ladder with frozen 8B extrapolation", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b and skywork-v2-qwen3-1.7b and skywork-v2-qwen3-4b and skywork-v2-qwen3-8b, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › >= 2 confirmed
  4. Susceptibility chi predicts best-of-n feature drift on a real bank Susceptibility chi predicts best-of-n feature drift on a real bank confirmed registered confidence interval 0.7500000000000002 to 1 > 0.3 The frozen study "Susceptibility chi predicts best-of-n feature drift on a real bank", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-8b and skywork-v2-llama31-8b, over ultrafeedback-bank. It does not establish that the same holds on another subject, another dataset or another run. split-bank estimator: chi and drift read disjoint prompt halves, so the null re-centers at zero, but the correlation's exchangeable unit is k = 6 features and the null still passes the frozen 0.3 threshold 0.28 of the time; the registered CI exclusion carries size control and only orderings near the synthetic calibration (0.958) clear the 0.91 MDE. Open the record › > 0.3 confirmed
  5. Certified erasure removes the exploit without benchmark loss Certified erasure removes the exploit without benchmark loss mixed registered > 0.8 The frozen study "Certified erasure removes the exploit without benchmark loss", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on hackfore-flagged, over hack-probes and rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.8 confirmed
  6. Does the process reward model verify or style-read Does the process reward model verify or style-read refuted registered > 0.7 The frozen study "Does the process reward model verify or style-read", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on qwen-prm, over processbench-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.7 refuted
  7. Verdict before critique on a real generative judge Verdict before critique on a real generative judge confirmed registered >= 0.9 The frozen study "Verdict before critique on a real generative judge", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-critic, over judge-pairs-1000. It does not establish that the same holds on another subject, another dataset or another run. Open the record › >= 0.9 confirmed
  8. Skepticism and receipt reliance predict who rewards fabricated receipts Skepticism and receipt reliance predict who rewards fabricated receipts refuted registered > 0.5 The frozen study "Skepticism and receipt reliance predict who rewards fabricated receipts", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet-subset, over receipt-triples-1600. It does not establish that the same holds on another subject, another dataset or another run. fraction of positive grader contrasts at n = 5 graders: the frozen 0.5 sits at the null center, the point call passes under the null 0.51 of the time, and 80 percent confirmation needs per-grader positive-contrast probability 0.67, the s15-plant regime (calibration 0.993). Open the record › > 0.5 refuted
  9. Value-convergence excess across the fleet beats the capability-matched null Value-convergence excess across the fleet beats the capability-matched null refuted registered confidence interval -0.12281662932807336 to -0.009192229234832329 > 0 The frozen study "Value-convergence excess across the fleet beats the capability-matched null", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-helpfulness-512. It does not establish that the same holds on another subject, another dataset or another run. the vce > 0 margin sits at the null center by construction: the point call passes under the null 0.48 of the time and its power at the threshold-alternative equals that rate; the alpha-exact p-value companion and the registered CI exclusion carry the card's evidence. Open the record › > 0 refuted
  10. The fleet decodes a property it does not price The fleet decodes a property it does not price refuted registered > 0.2 The frozen study "The fleet decodes a property it does not price", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.2 refuted
  11. Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor mixed registered > 0.6 The frozen study "Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet and armorm, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.6 refuted
  12. The RM knows it is being tested, and steering along that direction inflates reward The RM knows it is being tested, and steering along that direction inflates reward mixed registered > 0.6 The frozen study "The RM knows it is being tested, and steering along that direction inflates reward", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on evalaware-rms, over evalaware-paired and wildchat-organic and arena-organic. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0.6 confirmed
  13. Intransitive preference mass no scalar reward can represent Intransitive preference mass no scalar reward can represent confirmed registered The frozen study "Intransitive preference mass no scalar reward can represent", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on nectar-tournaments and ultrafeedback-tournaments. The intransitive-mass reading is dead: zero cyclic triples in 186,378 triangles, 10,000 of 10,000 sampled tournaments acyclic, and the measured value sits on the encoding floor. The store still carries the adjudication as confirmed, which is why this cannot be left to the outcome mapping: it would render as supported. Open the record › > 0.03 confirmed
  14. Surface biases enter the reward direction before quality features Surface biases enter the reward direction before quality features mixed registered > 1 The frozen study "Surface biases enter the reward direction before quality features", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b, over ultrafeedback-train-20k. It does not establish that the same holds on another subject, another dataset or another run. entry-ratio point call at the frozen > 1.0, the null median: power equals the null pass rate 0.501, ratios at or above 2.60 confirm at 80 percent, and the s07 plant sits far above that; passes only through the explicit accept-underpowered path. Open the record › > 1 refuted
  15. Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk refuted registered < 0.15 The frozen study "Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b, over caltransfer-organisms. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.15 refuted
  16. Attribution versus patching, adjudicated against planted ground truth Attribution versus patching, adjudicated against planted ground truth refuted registered confidence interval -0.7735849056603774 to -0.4276729559748428 > 0 The frozen study "Attribution versus patching, adjudicated against planted ground truth", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on skywork-v2-qwen3-0.6b and skywork-v02, over adjavp-organism and rb2-patch-40. It does not establish that the same holds on another subject, another dataset or another run. Open the record › > 0 refuted
  17. Do the indices carry information beyond size, accuracy, and lineage Do the indices carry information beyond size, accuracy, and lineage mixed registered < 0.05 The frozen study "Do the indices carry information beyond size, accuracy, and lineage", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full. It does not establish that the same holds on another subject, another dataset or another run. Open the record › < 0.05 confirmed

What this got wrong

The campaign reported a rank correlation of -0.256. The value measured from the fixture is -0.171, and it supersedes it. The old number is still live in the shipped README. It is written out here because the correction is the content, and this is the only route on the site where that string is allowed to appear.

-0.256
reward-lens-clean/docs/BUILD_NOTES.md:44
-0.171
reward-lens-clean/docs/BUILD_NOTES.md:50

The campaign scored its own calibration

It also asked whether its own directional predictions had been any good, and adjudicated that question by the same machinery it used on everything else. The answer, and the numbers behind it, are in the third act.

Is the measurement real

The cards

All 27 are here, grouped by verdict. The inconclusive ones are kept because removing them would be selection, and this is the page where selection would be least defensible.

How the adjudication worked