documentation
The manual
White-box tools and built-in epistemic discipline for reward models, the objects that define what RLHF optimizes.
pip install reward-lens uv add reward-lens Getting started
Three ways in. One is a browser tab and installs nothing.
Tutorials
Two ways in, and they answer different questions.
How-to guides
You know the job and you want the exact calls.
Models and signals
Will these instruments run on my grader?
Instruments
You have a reward model and a preference pair. What can you actually measure about the decision?
Training loops
Where does the reward model sit in the loop, and where do the instruments clip on?
Concepts
What do you have to believe before any of the tools make sense?
The measurement discipline
What does it take to trust a number about a reward model?
API reference
Which import surface owns a name, and will reaching for it pull in torch?
Usage: reward-lens [OPTIONS] COMMAND [ARGS]... Operator surface for the reward-lens kernel: cards, scoreboard, claims, and the Atlas. ╭─ Commands ───────────────────────────────────────────────────────────────────╮ │ card Build an RM Card for a signal: a view over every stored Evidence │ │ about it (section 2.15). │ │ scoreboard Print the theorem scoreboard: standing theorems and candidate │ │ laws (section 2.14). │ │ claims Check documents against the store; exit nonzero if any number is │ │ unbound (section 2.15.5). │ │ score Score inputs with a reward model (GPU-gated). │ │ serve Serve a reward model as an RL-loop-compatible endpoint │ │ (GPU-gated). │ │ audit Run the blind auditing game against a signal or organism │ │ (GPU-gated). │ │ study Freeze, run, and report frozen studies (gate 3). │ │ atlas The reward-model population Atlas. │ │ organism The ground-truth organism foundry. │ ╰──────────────────────────────────────────────────────────────────────────────╯
The measurement discipline these pages document was put through a pre-registered audit in July 2026: 27 cards, 53 frozen hypotheses, refutations published with the same machinery as confirmations. The scored ledger is at /campaign/.