RECORDED SCORE TREE measured

Reference GRPO run, one grader leaf, 1,600 rollouts

leaves
1
length_reward
0.62 to 3.36
abstained
200 of 1600

The score tree on this run has one leaf, named length_reward. 1,400 of 1,600 rollouts carry a value between 0.62 and 3.36; on the other 200 the grader abstained and the leaf holds nothing. One leaf is not a decomposition, and the spectrum here has one band for that reason. 101 of 200 drawn, uniform stride over the recorded values, first and last kept, no randomness.

reward-lens-assay/tests/fixtures/grpo_run/long/runs/run_f77bf75940ab982bbc35407af99cc094 · the run record's step, group and trajectory tables, in two shards of 100 steps

This run's score tree has exactly one leaf, so the spectrum here shows what a decomposition looks like when there is nothing to decompose. It establishes nothing about how a reward built from several components divides between them. On 200 of the 1,600 rollouts the grader abstained, and an abstention is an absent value rather than a low score.

VERIFIER AUDIT measured

MATH equivalence checker, MATH-500 gold answers, 240 rollouts, 60 mutations

killed
0 of 60
grader calls
2400
read, full access
5 of 13
read, training log only
0 of 13
metamorphic violations
259 of 378

full 5 of 13

log only 0 of 13

MATH equivalence checker: 0 of 60 competent mutants killed over 2,400 grader calls on MATH-500 gold answers, 240 rollouts. A reader with full access to this grader gets 5 of its 13 card fields; a reader with only the training log gets 0.

src/data/grader-cards.json · the audit card for is_equiv, with its full-access and log-only readings

Zero of 60 is an absence of detection on this one subject. It is not a population estimate and it is not a general defect rate for verifiers. Metamorphic violations are counted against this card's own denominator of 378 applicable pairs; a differently defined count of the same name exists elsewhere and is not mixed with this one.

VERIFIER AUDIT measured

SWE-bench test-report grader, SWE-bench Lite gold test lists, 240 rollouts, 60 mutations

killed
2 of 60
grader calls
2400
read, full access
5 of 13
read, training log only
0 of 13

full 5 of 13

log only 0 of 13

SWE-bench test-report grader: 2 of 60 competent mutants killed over 2,400 grader calls on SWE-bench Lite gold test lists, 240 rollouts. A reader with full access to this grader gets 5 of its 13 card fields; a reader with only the training log gets 0.

src/data/grader-cards.json · the audit card for swebench, with its full-access and log-only readings

2 of 60 is a reading on this one subject and this one corpus. It is not a population estimate. No metamorphic relation applied to this subject, so no violation bar is drawn.

VERIFIER AUDIT measured

GSM8K answer grader, GSM8K test split, 240 rollouts, 60 mutations

killed
24 of 60
grader calls
2400
read, full access
6 of 13
read, training log only
1 of 13
metamorphic violations
122 of 430

full 6 of 13

log only 1 of 13

GSM8K answer grader: 24 of 60 competent mutants killed over 2,400 grader calls on GSM8K test split, 240 rollouts. A reader with full access to this grader gets 6 of its 13 card fields; a reader with only the training log gets 1.

src/data/grader-cards.json · the audit card for verl_gsm8k, with its full-access and log-only readings

24 of 60 is a reading on this one subject and this one corpus. It is not a population estimate. Metamorphic violations are counted against this card's own denominator of 430 applicable pairs; a differently defined count of the same name exists elsewhere and is not mixed with this one.

VERIFIER AUDIT measured

Open-domain QA exact-match grader, NQ-Open validation split, 240 rollouts, 60 mutations

killed
26 of 60
grader calls
2400
read, full access
6 of 13
read, training log only
1 of 13
metamorphic violations
2 of 362

full 6 of 13

log only 1 of 13

Open-domain QA exact-match grader: 26 of 60 competent mutants killed over 2,400 grader calls on NQ-Open validation split, 240 rollouts. A reader with full access to this grader gets 6 of its 13 card fields; a reader with only the training log gets 1.

src/data/grader-cards.json · the audit card for verl_search_r1, with its full-access and log-only readings

26 of 60 is a reading on this one subject and this one corpus. It is not a population estimate. Metamorphic violations are counted against this card's own denominator of 362 applicable pairs; a differently defined count of the same name exists elsewhere and is not mixed with this one.

INTERVENTION FIXTURE measured

Skywork-Reward-Llama-3.1-8B-v0.2, 12 dimensions, 64 components, two localizers

dimensions
12
components
64
mean rank correlation
-0.170885
code correctness
-0.441117

Two ways of asking which part of a reward model produced a score, on the same 12 dimensions and the same 64 components. Their rank correlation averages -0.170885 across the dimensions, and on code correctness it reaches -0.441117. The largest single patching movement is 7.71 times the pair margin, and it is drawn at that size rather than clamped.

src/data/intervention.json · the per-component attribution and patching table, both ranks on every component of every dimension

The attributed set holds 65 components and the patched set holds 64; the embed is attributed and not patched, and the 64 drawn here are the ones both localizers cover. This is a disagreement measured on one reward model. On another model in the same fixture the same correlation is near zero, which is why it is not stated as a general result about attribution.

RECORDED SELECTION measured

Reference GRPO step, 8 rollouts in 2 groups, real advantages

rollouts
1600
advantage
-1.49161 to 1.493695
exactly zero
220
abstained
200

Each ray is one rollout and its offset is that rollout's recorded advantage, the value the estimator actually applied. Across 200 steps the 1,600 advantages run from -1.49161 to 1.493695. 220 of them are exactly zero, and 200 of those are rollouts the grader abstained on, which this estimator forces to zero rather than dropping. 128 of 1600 drawn, uniform stride over the recorded values, first and last kept, no randomness.

reward-lens-assay/tests/fixtures/grpo_run/long/runs/run_f77bf75940ab982bbc35407af99cc094 · the trajectory table's advantage column and the group table's mean and standard deviation

Nothing here is per token. The per-token advantage column is empty on all 1,600 rollouts and the recorded token count is zero throughout, so this record supports no per-position geometry and none is drawn. 220 rays sit on the axis because zero is what was recorded for them, which for 200 of them means the grader abstained rather than that the rollout was average.

RECORDED UPDATE measured

Reference GRPO step, 8 rollouts, immediate credit

steps
200
rollouts per step
8
clipped gradient norm
0.253293 to 0.441912

Across 200 recorded steps the clipped gradient norm runs from 0.253293 to 0.441912, on batches of 8 rollouts each. What the record does not hold is any per rollout gradient, so how much of that norm survived cancellation between the eight rollouts is not drawable from it. 101 of 200 drawn, uniform stride over the recorded values, first and last kept, no randomness.

reward-lens-assay/tests/fixtures/grpo_run/long/runs/run_f77bf75940ab982bbc35407af99cc094 · the step table and its optimizer payload

No per-rollout gradient is stored anywhere in this record, so the share of gradient mass that cancels within a step cannot be drawn from it. The project reports residual figures for immediate credit from a separate experiment; those exist as prose and as test assertions with no persisted array behind them, so they are stated in the text and are not encoded in any geometry here.

NO MEASUREMENT schematic

Policy response, nothing recorded

Nothing was measured here. The reconciliation this state would have illustrated does close on the 200-step run, and the closure is carried by the sampling noise of two batches of eight rollouts. On the worst feature-run pair the test's own detection floor is 3.760 against a first-order prediction of 5.647e-06, a ratio of 665,862; the median ratio over the six pairs is 854,374. Monte Carlo uncertainty is effectively all of the composed variance, and five of the nine uncertainty terms were computed while four are named absent. Neither verdict could have failed meaningfully, and that is the result this state shows.

reward-lens-assay/FINDINGS.md · the project's own findings record, which states the detection floor, declines the flattering verdict, and names the four budget terms it could not compute

The negative is the content. A response geometry drawn here would be an invention, and three times in this build's history a central visual claim turned out to have no measured instance behind it. The figures in the caption come from the project's own reconciliation report and are prose: they are stated, cited, and never encoded as a width, a count or a distance.

PUBLIC TRAINING SERIES measured

Public reward-hacking run, 25,664 labelled rollouts, one scalar-series warning result

labelled rollouts
25664
drawn points
101
transition midpoint
step 106
promoted signal alarms
step 90
fires on in-control surrogates
24.3%
detectors drawn
2 of 4
no alarm or rate recorded
2 of 4

25,664 labelled rollouts, 64 at each of 401 recorded step indices. The labelled series crosses its midpoint at step 106 with a ten-to-ninety width of 23.9 steps. The promoted signal alarms at step 90 and fires on 24.3% of in-control surrogates. The signal recorded as "Within-group reward variance, level" alarms earlier, at step 70, and fires on 60.7%, which is why it was rejected. The figure draws 401 points. The frozen ledger states 400 steps for this run and the published rollout table carries 401 distinct step indices; the site states the lower number and does not assert which record is right. 101 of 401 drawn, uniform stride over the recorded values, first and last kept, no randomness.

src/data/run-series.json · the per-step means of the published rollout table, with the fitted transition and the four evaluated detectors

Two records of this same run disagree on its length: the frozen ledger says 400 steps and the published rollout table carries 401 distinct step indices. The figure draws 401 points and the site does not assert which record is right. The earliest alarm here is not the best detector: the signal that fires first is the one the study rejected, at a surrogate false-alarm rate of 60.7%. One run does not establish a detection rate.

Follow the signal

The scorer

The reward went up. What changed inside the model?

Reward Lens instruments the path from a grader's score to the advantages training used, the update that followed, and the policy that came out. Inspect what was favoured, what moved, and what the account still cannot explain.

pip install reward-lens copy copied select to copy

installs 3.0.0

The published release is 3.0.0. This is what pip install reward-lens gets.
Follow the signal

One grader's comparison energy enters a dispersing element from the left and separates into labelled bands of different widths, ending at a detector plane. The widest band is the part that corresponds to a consistent ordering. The narrower one beside it is the part that does not. A third band is drawn at zero width, because that component was measured and found to be absent.

A score has to pass through a training process before it can change a model.

  1. The scorer

    The score is where the investigation starts, not where it ends.

    People specify an outcome. A real training system then executes something else on its behalf: a grader, a verifier, a reward model, or a program. The number it returns may be composed from several parts, sensitive to surface features, or reachable by a route nobody intended.

    Reward Lens keeps that scoring process inspectable wherever the available access allows, and preserves the components, gates and overrides before they collapse into one number.

  2. The effective training signal

    A score is not an update.

    In the policy-gradient runs shown here the optimizer never sees the score. It sees an advantage. Grouping, baselines, normalization, masks and clipping all sit between the two, and any of them can change the sign as well as the size of what a rollout contributes.

    Keeping the score tree and the run record intact makes the alternative computable: change one component of the reward and the same recorded rollouts can be scored again, without calling the grader, to see which of them would have been reinforced instead.

  3. The update

    One update has an address.

    The optimizer computes local credit and then discards the decomposition, keeping only the summed direction. Reward Lens partitions that immediate gradient back across the rollouts that produced it, and checks the partition against an independently computed full update rather than trusting it.

    Contributions align, oppose, and cancel. On the recorded step shown here more than half of the per-rollout gradient mass cancels before the optimizer sees any direction at all.

    This audits one step at the current parameters. It does not assign long-run responsibility for a behaviour that emerged over a whole run.

  4. The policy response

    The policy is the object that changed.

    Weight distance is not the question. Behaviour and internal features can move differently from each other, and both can move differently from what the training signal selected for. Four quantities are distinct here: the pressure applied, the response the policy could cheaply make, the movement actually observed, and the part of that movement nothing accounts for.

    This is the account Reward Lens is built to make close or visibly fail. On the run available here it does not close. The budget balances, but it balances on the sampling noise of two small batches, and the test's own detection floor sits hundreds of thousands of times above the effect it is arbitrating. Neither verdict could have failed meaningfully, and that is the honest reading.

  5. The live run

    A useful warning arrives before the decision is irreversible.

    On one public labelled run, a variance-derivative signal alarms before the fitted behavioural transition, and the named gradient-norm comparator peaks after it. The interesting part is not the lead. It is the control.

    Two signals appeared to lead by more and were both discarded, because on order-destroyed surrogates built from the same series they fire on roughly two thirds of cases. A signal that alarms on noise has not detected anything. Separating those was the work.

    One run, one scalar time series. No activations were read and no mechanism was named. An alarm and an explanation are different achievements.

What the instrument can already show

Three places where the ordinary training story loses information.

Where a reward appears is not where it is caused.

Rank a reward model's components by how much attribution assigns them, then rank them again by how much causal patching says they carry. On one widely used model the two orderings disagree, sharply on some objectives. On another they are almost unrelated, in the other direction. Attribution answers where the reward is visible. Patching asks what changes it when intervened on. Neither number licenses the other's claim.

subject
Skywork/Skywork-Reward-Llama-3.1-8B-v0.2
dimensionCount
12
attributedComponentCount
65
patchedComponentCount
64
unpatchedComponents
embed
patchingMode
noising

reward-lens-assay/fixtures/e_parity/golden.json:9

Fig. 1 · Attribution rank against patch-effect rank One model, twelve dimensions, one frozen study.

Twelve objective dimensions on one model. Above the zero line the two methods agree on a component's importance; below it they disagree. Eight of twelve fall below.

Inside a reward model

Change the reward. Watch the update change with it.

Open a reward score tree, disable one component, and see which of the recorded advantages change sign. Then follow a single rollout from its score, through its advantage, into its signed contribution to the update, and compare the summed partition against an independently computed full gradient.

Follow a training update

The policy is only as aligned as the pressure applied to it.

A grader card treats a scoring program as a measurement device rather than a leaderboard row: its source, its corpus, its branch coverage, how much of it a deliberate mutation can change without being caught, and which of its fields are even readable under the access you actually have. On one widely used equivalence checker, none of the tested mutations was caught.

subjectCount
4
corpus
MATH-500 gold answers, 240 rollouts
graderCalls
2400
fingerprint
verifier:95be1ad8dbf1474a1a7d56058e5f2b25
source
https://raw.githubusercontent.com/hendrycks/math/main/modeling/math_equivalence.py (sha256 c4101b1f51a2bb65)
Fig. 2 · Sixty mutations, and the answer key that noticed none of them
  • hendrycks/math is_equiv

    killedMutants 0

  • SWE-bench grading.py

    killedMutants 2

  • verl GSM8K reward

    killedMutants 24

  • verl Search-R1 QA exact match

    killedMutants 26

Each mark is one mutant of the equivalence checker, run against its own benchmark's gold answer pairs. A mutant is killed when the pairs disagree with the original. None were.

exploratory

One checker, and the comparison rests on two of the other three subjects rather than all three: the same instrument kills a substantial share of mutants on both verl environments and almost none on SWE-bench's report builder. That is what makes a score of zero a finding rather than a broken harness. The counts are in the table.

Audit a grader

Every claim says how far it reaches.

A result carries its subject, its access, its uncertainty, its comparator, its source and its scope. When the available record cannot support an answer, the result names what is missing and what would make the question answerable, rather than returning a worse number. A measured zero and a missing measurement are never allowed to look the same.

hendrycks/math is_equiv · full

coverage

read exploratory

Across 240 rollouts is_equiv has never taken 32% of its branches (12 of 38). Those are behaviours it cannot distinguish.

  • never fired: _fix_fracs.clause4, _fix_fracs.clause5, _remove_right_units.clause1, _strip_string.clause1, _strip_string.clause5, is_equiv.clause1, is_equiv.clause2, is_equiv.clause3
  • a random sample of the same size reaches 0.653 against the corpus's 0.684

hendrycks/math is_equiv · log_only

coverage

ACCESS_INSUFFICIENT declined

needs GRADER: SOURCE; you have GRADER: RECORD, RECORD: RECORD

remedy

supply grader at SOURCE. If you cannot, ask for the same quantity at a lower rung: `reward-lens capabilities` prints which rungs your access reaches and what each costs.

What this has earned, and what it has not.

Works now

  • Inspect reward models with observational and causal instruments.
  • Preserve composed scores and replay them against the advantages a run actually used.
  • Audit where one immediate update came from, against an independent full gradient.
  • Inspect graders and verifier programs, including what a mutation can change unnoticed.
  • Measure behavioural and representational movement in stated feature spaces.
  • Return a sourced reading, or a specific reason it cannot read.

Not established

  • A general decoder of a policy's true objective.
  • A causal explanation of a naturalistic reward-hacking run.
  • Forecasting that holds across models, seeds and mechanisms.
  • Separating a target from a correlated passenger, in practice.
  • Long-horizon credit across many turns and tool calls.
  • A closed account of selection through to policy response, at an informative scale.

Everything this does not do

Start with your run.

pip install reward-lens copy copied select to copy

installs 3.0.0

The published release is 3.0.0. This is what pip install reward-lens gets.
bash
reward-lens capabilities

The capability report reads what your record and your access can actually support. It names which questions are answerable, what each one costs, and which information was never recorded, before you spend anything.