See what the reward model actually learned.

White-box tools and built-in epistemic discipline for reward models, the objects that define what RLHF optimizes.

pip install reward-lens
uv add reward-lens
reward-lens --help
 
 Usage: reward-lens [OPTIONS] COMMAND [ARGS]...

 Operator surface for the reward-lens kernel: cards, scoreboard, claims, and
 the Atlas.
╭─ Commands ───────────────────────────────────────────────────────────────────╮
│ card        Build an RM Card for a signal: a view over every stored Evidence │
│             about it (section 2.15).                                         │
│ scoreboard  Print the theorem scoreboard: standing theorems and candidate    │
│             laws (section 2.14).                                             │
│ claims      Check documents against the store; exit nonzero if any number is │
│             unbound (section 2.15.5).                                        │
│ score       Score inputs with a reward model (GPU-gated).                    │
│ serve       Serve a reward model as an RL-loop-compatible endpoint           │
│             (GPU-gated).                                                     │
│ audit       Run the blind auditing game against a signal or organism         │
│             (GPU-gated).                                                     │
│ study       Freeze, run, and report frozen studies (gate 3).                 │
│ atlas       The reward-model population Atlas.                               │
│ organism    The ground-truth organism foundry.                               │
╰──────────────────────────────────────────────────────────────────────────────╯
The 53 adjudicated hypotheses: card, hypothesis, registered prediction, observed value, outcome.
CardHypothesisRegisteredObservedOutcome
ERR-RB2H-err-qwenerr_auroc_delta_qwen8b > 0.03-9.74e-4refuted, K-err fired
ERR-RB2H-err-llamaerr_auroc_delta_llama8b > 0.03-0.003471refuted, K-err fired
FORECAST-CATH-forecast-maeforecast_mae < 0.060.08696refuted
FORECAST-CATH-forecast-beats-sizeforecast_mae_minus_size_baseline < 0-0.1024confirmed
STYLE-RMBH-style-transferspearman_biasbattery_vs_rmbench_hard > 0.6not measuredinconclusive
STYLE-RMBH-style-permbiasbattery_rmbench_perm_p < 0.05not measuredinconclusive
STYLE-RMBH-style-baselinespearman_minus_behavioral_baseline >= 0not measuredinconclusive
PPE-BONH-tail-plateauspearman_tail_vs_bon_plateau > 0.3not measuredinconclusive
PPE-BONH-lowerquantilebon_lowerquantile_beats_mean == 1not measuredinconclusive
LADDERH-ladder-coverageladder_8b_metrics_in_interval >= 23confirmed
LADDERH-ladder-beats-flatladder_extrap_minus_flat_error < 00.6002refuted
CHI-DRIFTH-chi-bonchi_bon_spearman > 0.30.9643confirmed
CHI-DRIFTH-chi-beats-varchi_minus_variance_baseline_spearman > 00.5714confirmed
HUMPH-hump-existshump_interior_present == 1not measuredinconclusive
HUMPH-hump-lochump_kl_abs_error < 0.5not measuredinconclusive
GAUGE-E19H-rawraw_cos_v01_v02 abs< 0.02not measuredinconclusive
GAUGE-E19H-canoncanonical_minus_raw > 0.4not measuredinconclusive
GAUGE-E19H-localizeresidual_concept_matches_top_behavioral_delta == 1not measuredinconclusive
GAUGE-XFAMH-xfamspearman_canonical_angle_vs_disagreement > 0.6not measuredinconclusive
GAUGE-XFAMH-xfam-permxfam_perm_p < 0.05not measuredinconclusive
HACK-FOREH-flagflag_hit_rate_minus_behavioral > 0not measuredinconclusive
SURGERYH-eraseexploit_drift_reduction > 0.80.8856confirmed
SURGERYH-retainrb2_accuracy_delta_after_erasure > -0.01-0.3988refuted
VERIF-PRMH-verif-locdense_localization_auc > 0.70.2821refuted, K-verif fired
VERIF-PRMH-verif-stylestyle_share < 0.5-0.4268confirmed, K-verif fired
JUDGE-VBCH-vbcverdict_prefix_match_rate >= 0.91confirmed
VALUES-CONTESTH-contestcontested_raterspread_spearman > 0.2not measuredinconclusive
VALUES-CONTESTH-contest-vs-ensembleensemble_minus_contested_spearman < 0.05not measuredinconclusive
FORENSIC-RECEIPTH-reliancereliance_recovery > 0.50refuted, K-forensic fired
FORENSIC-RECEIPTH-credulityomission_credulity_gap > 0.20.02326refuted, K-forensic fired
ATLAS-VCEH-vcereal_vce > 0-0.06987refuted, K-vce fired
ATLAS-VCEH-vce-preward_convergent_p_value < 0.050.001confirmed, K-vce fired
FACT-KUIH-kui-gapkui_gap_fleet_median > 0.20.1515refuted, K-kui fired
FACT-KUIH-kui-permkui_gap_perm_p < 0.050.3886refuted, K-kui fired
FACT-KUIH-kui-controlkui_control_fleet_median < 0.10.3536refuted, K-kui fired
CAPACITY-WELCHH-dark-kdeffdark_reward_kdeff_spearman > 0.6-0.4062refuted
CAPACITY-WELCHH-capacity-permcapacity_perm_p < 0.050.8934refuted
CAPACITY-WELCHH-welchwelch_floor_min_slack >= 00.8164confirmed
EVAL-AWAREH-aware-probereal_probe_balanced_acc > 0.60.7932confirmed
EVAL-AWAREH-aware-steerdelta_r_per_steer > 0.05-2.10e-4refuted
EVAL-AWAREH-aware-organic-fprorganic_fpr < 0.10.317refuted
TOPO-HODGEH-intransitiveintransitive_mass > 0.030.214confirmed
EMB-LORAH-bias-firstreal_bias_before_quality > 10.2692refuted
EMB-LORAH-stabilizew_r_stabilization_monotone == 11confirmed
CAL-TRANSFERH-transfermax_abs_auc_difference < 0.150.419refuted, K-transfer fired
ADJ-AVPH-avprecovery_gap > 0-0.566refuted, K-avp fired
CONF-PARTIALH-partial-permbest_index_perm_p < 0.050.0468confirmed
CONF-PARTIALH-partialmax_abs_partial_corr > 0.30.6485confirmed
CONF-PARTIALH-partial-lineagelineage_check_pass == 10refuted
META-LEDGERH-brierbrier_directional < 0.250.26refuted, K-meta fired
META-LEDGERH-coverageinterval_coverage_0p8 >= 0.60.75confirmed, K-meta fired
T3-DECOMPH-tacitreal_tacit_fraction > 00.9723confirmed
T3-FIELDH-flatflat_hack_overlap_minus_random > 0not measuredinconclusive

53 pre-registered hypotheses about ten reward models, adjudicated July 2026: 16 confirmed, 21 refuted, 16 inconclusive, 8 kills fired. Hover any mark.

One output direction

A reward model reads text, forms a hidden state h, and one learned direction turns it into the score: r(x) = w_r · h + b. Everything on this site is about that projection and what feeds it.

The strip here is real, precomputed data: per-token contributions to the score of the smallest fleet model, computed with the library's dense instrument on its built-in diagnostic pairs. On the formatting pair the model prefers the rejected completion; that miss ships as is.

The reward projection: a hidden state read out along one learned direction

prompt What's 15% of 80?

<|im_start|>userWhat's 15% of 80?<|im_end|><|im_start|>assistant<think>⏎⏎</think>⏎⏎15% of 80 is 12.<|im_end|>

score gap: the model prefers the chosen completion by 5.65

Skywork/Skywork-Reward-V2-Qwen3-0.6B · signals.dense.dense_rewards (first differences of the prefix-score curve) · trust E by construction: dense maps stay exploratory until certified against labeled error spans

Per-token reward contributions for the chosen completion of the helpfulness pair.
TokenContribution
<|im_start|>0.733
user2.76
-1.46
What-3.49
's1.12
-3.68
13.51
51.93
%-2.14
of2.08
-3.82
81.49
01.99
?1.29
<|im_end|>-3.58
1.88
<|im_start|>-1.66
assistant4.63
-3.61
<think>-4.37
-0.543
</think>3.29
1.89
1-7.24
55.29
%0.916
of-2.11
-5.58
84.8
01.46
is3.77
-9.55
15.89
23.03
.-1.95
<|im_end|>3.54
-0.235

The instruments

Attribute a score to heads and layers

mb.run(DirectLinearAttribution(), mb.Context(signal=signal, view=view))

Patch and verify causally

mb.run(PatchGrid(), mb.Context(signal=signal, view=view))

Run the bias battery

mb.run(BiasBattery(), mb.Context(signal=signal, view=view))

Track w_r formation over checkpoints

from reward_lens.dynamics import CheckpointSequence

Forecast drift under best-of-n

from reward_lens.loops import SusceptibilitySpectrum, BoNLadder

Freeze a study and adjudicate it

reward-lens study freeze spec.json

The first pre-registered white-box audit of reward models

In July 2026 the discipline was put through a live fire exercise: 27 study cards, 53 hypotheses with frozen thresholds and kill criteria, predictions hashed before the run, then adjudicated by machinery against a merged evidence store. 16 confirmed, 21 refuted, 16 inconclusive; 8 kill criteria fired, including the campaign's kill on its own calibration.

One row per headline call: position shows the margin by which the frozen call held or failed, color the adjudication
One row per headline call: right of the zero line means the frozen call held (less-than calls are sign-flipped, p-valued calls plot log10(threshold/p), equality calls can only ever land on the line itself, and interval rows shade (0, 1], the entire held region: 0 is the frozen band's edge, 1 is its dead center); colors carry the final adjudication, so a positive-margin marker downgraded by its CI or FDR overlay stays refuted with the drawn interval showing why; 7 held, 10 refuted, 0 inconclusive, counted from the rows drawn.

Refutations and missing evidence are published with the same machinery.

Read the scored ledger

Every number wears its trust level

E

exploratory

A bare measurement. Nothing has earned it more than that.

C

calibrated

The instrument holds a scorecard from runs against planted ground truth.

R

registered

The prediction predated the data, frozen and hashed.

A

adjudicated

Machinery compared it to its frozen threshold and wrote the verdict.

The gates compute these levels; no caller can override them.

Open problems, stated as measurable questions

Forecast the hack before any RL

Can the KL location of the Goodhart hump be predicted from internals alone? The drift ordering held; the hump location is unrun.

Calibration that transfers across scales

The 0.419 gap against a frozen 0.15 bound is now the problem statement.

Localize process errors in verifiers

The dense extractor failed its answer-key validation on a real PRM, so the instrument itself is open.

The science page carries the rest.