Three questions, three instruments, and what each one could not answer.
One reward model, one training run, one grader. Each one ends with the code, the data and the expected output, so you can run it yourself.
A policy is trained against a number. The number comes from a grader or a reward model, and the training run turns it into behaviour. Three things can therefore be looked inside: the model that holds the preference, the run that acts on it, and the program that produces the score. Each of the three below opens one of them.
They are three separate studies on three separate subjects, and no subject here carries the whole chain. A static reward model, a public labelled training run, and four scoring programs were each measured on their own. Reading all three tells you what each instrument does. It does not give you one system followed from its grader to its trained policy, and nothing on this site does yet.
This model prefers one answer over another. Where does that preference live, and if I intervene there, does the score move?
Skywork-Reward-Llama-3.1-8B-v0.2, twelve objective dimensions, forward and backward access to the model. Components ranked twice: once by observational attribution, once by causal patching.
Observational localization and causal localization can point in different directions on the same model and the same objective. A component that a visible representation makes prominent is a hypothesis about mechanism. It is not a causal result until something is patched.
Spearman rank correlation reward-lens-assay/fixtures/e_parity/golden.json:9 across twelve dimensions on this model. Eight of the twelve dimensions are negative and four are positive.
This does not establish that attribution is wrong and patching is right. On a model organism with a planted answer key, the prediction that patching would recover the key better than attribution was frozen before the data, and the result contradicted it. The recorded conclusion is that neither localization can be preferred. A noisy answer key is the disclosed alternative reading and it has not been ruled out.
My policy is training right now. Is the objective it is actually optimizing drifting away from the one I wrote, and will I know before it is obvious?
A public labelled reinforcement-learning run with reward hacking: reward-lens-assay/PREDICTIONS.md:154 labelled rollouts across reward-lens-assay/PREDICTIONS.md:154 training steps, or reward-lens-assay/PREDICTIONS.md:154 by the run's other record, which is the count its own rollout table carries. The hacking transition is fitted at step reward-lens-assay/PREDICTIONS.md:154 with a 10-to-90 width of reward-lens-assay/PREDICTIONS.md:154 steps.
The variance derivative alarms at step reward-lens-assay/PREDICTIONS.md:154, a lead of reward-lens-assay/PREDICTIONS.md:154 transition widths. The gradient-norm comparator, named in advance, peaks at step reward-lens-assay/PREDICTIONS.md:154, which is reward-lens-assay/PREDICTIONS.md:154 steps after the midpoint, a lead of reward-lens-assay/PREDICTIONS.md:154 widths.
The variance level alarms at step reward-lens-assay/PREDICTIONS.md:154 for a lead of reward-lens-assay/PREDICTIONS.md:154 widths and fires on reward-lens-assay/PREDICTIONS.md:154 of order-destroyed surrogates from the same series. The gradient norm under a cumulative-sum detector fires on reward-lens-assay/PREDICTIONS.md:154. The shipped derivative fires on reward-lens-assay/PREDICTIONS.md:154.
This is one scalar series on one public run. No activations were read and no mechanism is named. The false-positive rate across runs is not settled, and the gradient-norm comparator could not be computed on the companion analysis at all, because the published rollout table carries no optimiser telemetry.
Everything above depends on a grader producing scores. Is the grader sound?
Four scoring programs in wide use: a mathematical equivalence checker from a standard maths benchmark, the report builder inside a software-engineering benchmark's grading module, and two environment reward functions that ship with a widely used trainer.
reward-lens--github/reward-lens-assay/experiments/x1_release/cards/is_equiv.txt fields rendered, reward-lens--github/reward-lens-assay/experiments/x1_release/cards/is_equiv.txt read, reward-lens--github/reward-lens-assay/experiments/x1_release/cards/is_equiv.txt refused, and none of them for want of access.
A score of zero is an absence of detection on one checker against one corpus, at the denominator stated beside it. It is not a defect rate for verifiers in general and it is not an estimate of any population.
None of the four subjects is a rollout from a policy under training. This says what a grader does to inputs, not what it does to an optimisation run.