Settled by one campaign

The fluctuation-dissipation prediction for reward hacking has its first real-model confirmation: susceptibility measured on a base policy predicted realized best-of-n feature drift at rho chi_bon_spearman Value 0.964286 Subject CHI-DRIFT Threshold > 0.3 Trust Adjudicated Evidence ev:c978e3958769ecf93dcf079771a3aedb on one model and 1.0 on a second (standing theorem T9, via CHI-DRIFT). Value-convergence excess has its first real refutation at fleet size ten: real_vce Value -0.0698739 Subject ATLAS-VCE Threshold > 0 Trust Adjudicated Evidence ev:0c80d9086ebb0ad9704b88a8eefb01fa against the capability-matched null, so the apparent convergence of machine values was fully explained by shared capability structure (T13, refuted via ATLAS-VCE).

A real generative judge commits to its verdict before writing its critique: the verdict decoded at the pre-critique position matched the final verdict at rate verdict_prefix_match_rate Subject JUDGE-VBC Threshold >= 0.9 Trust Adjudicated Evidence ev:e7fd3fee89a0c56cf4465212601d0d26 (JUDGE-VBC). A fifth of the preference data's structure is cyclic mass no scalar head can represent: intransitive_mass Value 0.213976 Subject TOPO-HODGE Threshold > 0.03 Trust Adjudicated Evidence ev:844ad06a5aeb1d89c21c13c19a28cbd2 (TOPO-HODGE). And the meta-result: the campaign scored its own pre-run confidence and published the miss, a Brier of brier_directional Subject META-LEDGER Threshold < 0.25 Trust Adjudicated Evidence ev:3990df7af1488c08ad62786d879b0b0f against the coin's 0.25 (META-LEDGER, K-meta fired).

The standing board

Row Title Kind Status Science Adjudicated by
T1 Constructive unhackable-subspace finder standing ○ open S4
T2 Distortion equilibrium standing ○ open S8/S12
T3 RLHF speed proportional to teacher variance standing ○ open S12/S3
T4 Proxy-true reward angle standing ○ open S2/S12
T5 Heavy tail defeats KL control standing ○ open S3/S4
T6 Identifiability up to shift and scale standing ○ open S2
T7 No single scalar for a population standing ○ open S11
T8 Scalar head cannot express intransitivity standing ○ open S2
T9 Fluctuation-dissipation for reward hacking candidate law Confirmed S3 CHI-DRIFT
T10 Belief factorization and gauge=channel-kernel candidate law ○ open S8/S2
T11 Evaluator-model divergence precedes hacking candidate law ○ open S13/AT
T12 Coherence/Welch law and Hodge obstruction candidate law ○ open S5/S6
T13 Value convergence excess candidate law Refuted AT ATLAS-VCE
T14 Honesty unraveling law candidate law ○ open S15

The open problems

  1. Forecast the hack before any RL

    S12 first data point CHI-DRIFT

    Can the KL location of the Goodhart hump and the first-hacked dimension be predicted from internals before any RL? The drift ordering held on a real bank; the hump-location arm never ran.

  2. Calibration that transfers across scales

    S1 refuting data point CAL-TRANSFER

    Instrument scorecards moved from a CPU toy trunk to a real 0.6B trunk with a max per-instrument AUC gap of 0.419 against a frozen 0.15 bound. The gap is now the problem statement.

  3. Localize process errors in verifiers

    S9 instrument open VERIF-PRM

    Is the error represented at step k but priced at step k+2? The dense extractor failed its answer-key validation on a real PRM, so the instrument itself is open.

  4. Which part of a reward direction is physical

    S2 open GAUGE-E19

    Two RMs can be behaviorally identical with orthogonal reward directions. Both gauge cards await a clean run.

  5. Is learned preference rank-one

    S2/S6 first data point TOPO-HODGE

    The preference data carries 21 percent cyclic mass no scalar head can express. What does the head do with it?

  6. Hysteresis and phase structure of hacking

    S14 open

    Is the reward-hacking transition reversible? Does it sharpen with scale, like a genuine phase transition, or not?

  7. Does legibility grow or shrink with scale

    S10 first data point T3-DECOMP

    Tacit fraction 0.97 at 8B is one point on the curve. Nobody knows even the sign of its slope.

  8. Does the grader read the receipts

    S15 refuted on contact FORENSIC-RECEIPT

    The first receipt instruments were refuted on contact with real graders and are being rebuilt.

  9. The weak-to-strong elicitation boundary

    S13 open

    Above what salience does a weak grader elicit the strong model's own values rather than install its errors?

  10. Capacity floors

    S5 first data point CAPACITY-WELCH

    Is some bias information-theoretically forced once the rubric outgrows the reward-relevant subspace? The Welch floor held once.

  11. Kinship

    S13 open

    Does a policy hack its sibling grader faster than a foreign one at matched accuracy?

  12. What mediates eval-awareness

    S16 first data point EVAL-AWARE

    A probe reads benchmark-versus-organic at 0.79 balanced accuracy; steering the same direction moved reward by nothing. What carries the signal?

The sixteen sciences

One kernel, sixteen sciences, three gates.

S1 Ground truth and calibration
When attribution and patching disagree, which is right? Can a blind team handed reward-lens find a planted bias?
S2 Gauge and the scalar bottleneck
Is a near-zero cosine between two reward directions a functional change or a coordinate change? Is learned preference empirically rank-1?
S3 Reward thermodynamics
Are the feature-level consequences of optimizing against a grader derivable from base-policy statistics before any RL?
S4 Reward field theory
Are the flat directions of the reward Hessian the hackable subspace? Is the exploit surface finite, or a hydra?
S5 Capacity theory of bias
Does a scalar grader aggregating K criteria through a d-dimensional bottleneck suffer an irreducible interference floor?
S6 Preference topology
Is a computable fraction of reward error topologically obligatory, forced by intransitivity in the preference data?
S7 Reward embryology and pretraining origins
Does the reward direction form gradually or in phase transitions? Do surface biases enter before quality features?
S8 The knowledge-reward gap
Does the RM represent a property its reward ignores, the mechanistic precondition of hacking?
S9 Do verifiers verify
What fraction of an RM's correctness preference is causally anchored at the actual error versus carried by style?
S10 Decompiling and the legibility frontier
How much of the learned decision function can be said in language, and what is the tacit remainder?
S11 Values, pluralism and paradigm physiology
Do RMs encode contested pairs orthogonally to the reward direction? Is a judge's verdict decided before its critique is written?
S12 Hackability forecasting and certified prevention
Can the KL location of the Goodhart hump and the first-hacked dimension be predicted from internals before any RL?
S13 The recorder, coupling, robustness, kinship and weak-to-strong
Is hacking onset visible in reward-feature space before reward and KL curves move? When a weak grader teaches a strong model, which does the strong model learn: the values or the errors?
S14 Thermodynamic phase structure
What is the order of the reward-hacking transition? Does it exhibit hysteresis? Does it sharpen with scale?
S15 Trajectory forensics and the economics of grading
When a trajectory RM assigns reward, is it reading the receipts or the narrative?
S16 Certified robustness, attack surface and eval-awareness
After erasing an attack's mediating subspace, what token budget re-breaks the RM? Does the RM internally recognize a benchmark item?

Two cross-cutting meta-studies: universality of machine values (the VCE question, first refutation on the fleet) and performative interpretability (the audit half-life of each metric under developer response).

Collaboration runs through GitHub issues.