the science
What one campaign settled, and what it opened
Settled by one campaign
The fluctuation-dissipation prediction for reward hacking has its
first real-model confirmation: susceptibility measured on a base
policy predicted realized best-of-n feature drift at rho
chi_bon_spearman Value 0.964286 Subject CHI-DRIFT Threshold > 0.3 Trust Adjudicated Evidence ev:c978e3958769ecf93dcf079771a3aedb on one model and 1.0 on a second (standing theorem T9, via
CHI-DRIFT). Value-convergence
excess has its first real refutation at fleet size ten:
real_vce Value -0.0698739 Subject ATLAS-VCE Threshold > 0 Trust Adjudicated Evidence ev:0c80d9086ebb0ad9704b88a8eefb01fa against the capability-matched null, so the apparent convergence
of machine values was fully explained by shared capability structure
(T13, refuted via ATLAS-VCE).
A real generative judge commits to its verdict before writing its
critique: the verdict decoded at the pre-critique position matched the
final verdict at rate
verdict_prefix_match_rate Subject JUDGE-VBC Threshold >= 0.9 Trust Adjudicated Evidence ev:e7fd3fee89a0c56cf4465212601d0d26 (JUDGE-VBC). A fifth of the
preference data's structure is cyclic mass no scalar head can
represent:
intransitive_mass Value 0.213976 Subject TOPO-HODGE Threshold > 0.03 Trust Adjudicated Evidence ev:844ad06a5aeb1d89c21c13c19a28cbd2 (TOPO-HODGE). And the
meta-result: the campaign scored its own pre-run confidence and
published the miss, a Brier of
brier_directional Subject META-LEDGER Threshold < 0.25 Trust Adjudicated Evidence ev:3990df7af1488c08ad62786d879b0b0f against the coin's 0.25
(META-LEDGER, K-meta fired).
The standing board
| Row | Title | Kind | Status | Science | Adjudicated by |
|---|---|---|---|---|---|
T1 | Constructive unhackable-subspace finder | standing | ○ open | S4 | |
T2 | Distortion equilibrium | standing | ○ open | S8/S12 | |
T3 | RLHF speed proportional to teacher variance | standing | ○ open | S12/S3 | |
T4 | Proxy-true reward angle | standing | ○ open | S2/S12 | |
T5 | Heavy tail defeats KL control | standing | ○ open | S3/S4 | |
T6 | Identifiability up to shift and scale | standing | ○ open | S2 | |
T7 | No single scalar for a population | standing | ○ open | S11 | |
T8 | Scalar head cannot express intransitivity | standing | ○ open | S2 | |
T9 | Fluctuation-dissipation for reward hacking | candidate law | Confirmed | S3 | CHI-DRIFT |
T10 | Belief factorization and gauge=channel-kernel | candidate law | ○ open | S8/S2 | |
T11 | Evaluator-model divergence precedes hacking | candidate law | ○ open | S13/AT | |
T12 | Coherence/Welch law and Hodge obstruction | candidate law | ○ open | S5/S6 | |
T13 | Value convergence excess | candidate law | Refuted | AT | ATLAS-VCE |
T14 | Honesty unraveling law | candidate law | ○ open | S15 |
The open problems
-
Forecast the hack before any RL
Can the KL location of the Goodhart hump and the first-hacked dimension be predicted from internals before any RL? The drift ordering held on a real bank; the hump-location arm never ran.
-
Calibration that transfers across scales
Instrument scorecards moved from a CPU toy trunk to a real 0.6B trunk with a max per-instrument AUC gap of 0.419 against a frozen 0.15 bound. The gap is now the problem statement.
-
Localize process errors in verifiers
Is the error represented at step k but priced at step k+2? The dense extractor failed its answer-key validation on a real PRM, so the instrument itself is open.
-
Which part of a reward direction is physical
Two RMs can be behaviorally identical with orthogonal reward directions. Both gauge cards await a clean run.
-
Is learned preference rank-one
The preference data carries 21 percent cyclic mass no scalar head can express. What does the head do with it?
-
Hysteresis and phase structure of hacking
Is the reward-hacking transition reversible? Does it sharpen with scale, like a genuine phase transition, or not?
-
Does legibility grow or shrink with scale
Tacit fraction 0.97 at 8B is one point on the curve. Nobody knows even the sign of its slope.
-
Does the grader read the receipts
The first receipt instruments were refuted on contact with real graders and are being rebuilt.
-
The weak-to-strong elicitation boundary
Above what salience does a weak grader elicit the strong model's own values rather than install its errors?
-
Capacity floors
Is some bias information-theoretically forced once the rubric outgrows the reward-relevant subspace? The Welch floor held once.
-
Kinship
Does a policy hack its sibling grader faster than a foreign one at matched accuracy?
-
What mediates eval-awareness
A probe reads benchmark-versus-organic at 0.79 balanced accuracy; steering the same direction moved reward by nothing. What carries the signal?
The sixteen sciences
One kernel, sixteen sciences, three gates.
-
S1Ground truth and calibration - When attribution and patching disagree, which is right? Can a blind team handed reward-lens find a planted bias?
-
S2Gauge and the scalar bottleneck - Is a near-zero cosine between two reward directions a functional change or a coordinate change? Is learned preference empirically rank-1?
-
S3Reward thermodynamics - Are the feature-level consequences of optimizing against a grader derivable from base-policy statistics before any RL?
-
S4Reward field theory - Are the flat directions of the reward Hessian the hackable subspace? Is the exploit surface finite, or a hydra?
-
S5Capacity theory of bias - Does a scalar grader aggregating K criteria through a d-dimensional bottleneck suffer an irreducible interference floor?
-
S6Preference topology - Is a computable fraction of reward error topologically obligatory, forced by intransitivity in the preference data?
-
S7Reward embryology and pretraining origins - Does the reward direction form gradually or in phase transitions? Do surface biases enter before quality features?
-
S8The knowledge-reward gap - Does the RM represent a property its reward ignores, the mechanistic precondition of hacking?
-
S9Do verifiers verify - What fraction of an RM's correctness preference is causally anchored at the actual error versus carried by style?
-
S10Decompiling and the legibility frontier - How much of the learned decision function can be said in language, and what is the tacit remainder?
-
S11Values, pluralism and paradigm physiology - Do RMs encode contested pairs orthogonally to the reward direction? Is a judge's verdict decided before its critique is written?
-
S12Hackability forecasting and certified prevention - Can the KL location of the Goodhart hump and the first-hacked dimension be predicted from internals before any RL?
-
S13The recorder, coupling, robustness, kinship and weak-to-strong - Is hacking onset visible in reward-feature space before reward and KL curves move? When a weak grader teaches a strong model, which does the strong model learn: the values or the errors?
-
S14Thermodynamic phase structure - What is the order of the reward-hacking transition? Does it exhibit hysteresis? Does it sharpen with scale?
-
S15Trajectory forensics and the economics of grading - When a trajectory RM assigns reward, is it reading the receipts or the narrative?
-
S16Certified robustness, attack surface and eval-awareness - After erasing an attack's mediating subspace, what token budget re-breaks the RM? Does the RM internally recognize a benchmark item?
Collaboration runs through GitHub issues.