Instrument atlas
What can actually be measured, and what it costs to measure it.
An instrument measures one quantity, needs a stated level of access to do it, and returns either a reading with its trust level or a refusal naming what was missing. This catalogue is generated from the library's own registry rather than maintained by hand, so it cannot list an instrument the library does not have.
Find an instrument
Specified and not built. Listed here at the same weight as the rest, because a catalogue that quietly separates what exists from what is planned is one a reader stops trusting the moment they notice.
| What it measures | What access it needs | Implemented or open |
|---|---|---|
| Effective group size How many independent gradient-relevant observations a group of K rollouts actually gives you. | What access it needs: grader, record, replicate What kind of object: composite, human, neural_gen, neural_scalar, procedural, program When it can run: post_run, pre_run Free Signal validity: the grader as a measurement device | implemented |
| Variance components and gauge R&R Facet decomposition of `Var(r)`. | What access it needs: grader, replicate What kind of object: composite, human, neural_gen, neural_scalar, procedural, program When it can run: pre_run Signal validity: the grader as a measurement device | implemented |
| Attenuation factor How much grader error shrinks the selection signal. | What access it needs: grader, replicate Signal validity: the grader as a measurement device | implemented |
| Blackwell order and deficiency Rank graders by informativeness rather than accuracy. | What access it needs: grader, query Signal validity: the grader as a measurement device | implemented |
| Decision study: optimal allocation Where the next dollar goes. | What access it needs: not stated Signal validity: the grader as a measurement device | implemented |
| Grader stochasticity profile The grader as a distribution, not a number. | What access it needs: grader, replicate Signal validity: the grader as a measurement device | implemented |
| Environment flakiness Score variance attributable to the environment, not the model. | What access it needs: query, task What kind of object: program When it can run: pre_run Signal validity: the grader as a measurement device | implemented |
| Curl mass and the harmonic split How much preference structure no scalar can carry, and how much of that is fixable. | What access it needs: grader, query What kind of object: neural_gen, procedural When it can run: pre_run Grader structure: is this thing a scalar, and what is it made of? | implemented |
| Afriat efficiency index Rationalizability, from revealed-preference theory. | What access it needs: grader, record What kind of object: neural_gen, procedural Grader structure: is this thing a scalar, and what is it made of? | implemented |
| Composition tree and counterfactual composition What the reward is actually made of, and what it would be without a piece. | What access it needs: estimator, grader, record What kind of object: composite When it can run: post_run Grader structure: is this thing a scalar, and what is it made of? | implemented |
| Silent-zero and abstention rate How often the grader failed and returned a real number anyway. | What access it needs: grader, record What kind of object: composite, human, neural_gen, neural_scalar, procedural, program Grader structure: is this thing a scalar, and what is it made of? | implemented |
| Scalar-representability bound The worst-case number that carries the convergence theorem. | What access it needs: grader, query Grader structure: is this thing a scalar, and what is it made of? | implemented |
| Tournament solution and comparison-graph connectivity | What access it needs: grader, record What kind of object: procedural Grader structure: is this thing a scalar, and what is it made of? | implemented |
| The susceptibility triple `S`, `β`, and `Gβ`, always together. | What access it needs: backward, grader, policy, record Grader white-box: the existing battery, re-typed | implemented |
| Heritability and autonomy Which features can move at all, and what moves with them. | What access it needs: backward, policy, record Grader white-box: the existing battery, re-typed | implemented |
| Instrument recovery AUC The head-to-head nobody publishes. | What access it needs: mutate, organism Grader white-box: the existing battery, re-typed | implemented |
| Erasure, with the head-to-head that is now mandatory | What access it needs: grader, mutate Grader white-box: the existing battery, re-typed | implemented |
| Acute versus chronic intervention Ablate and keep training. | What access it needs: control, mutate, policy Grader white-box: the existing battery, re-typed | implemented |
| Rescue Put it back and check the behaviour returns. | What access it needs: mutate, policy Grader white-box: the existing battery, re-typed | implemented |
| Double dissociation Necessity is not sufficiency. | What access it needs: mutate, policy Grader white-box: the existing battery, re-typed | implemented |
| Jacobian-corrected verdict direction | What access it needs: forward, grader What kind of object: neural_gen When it can run: pre_run Grader white-box: the existing battery, re-typed | implemented |
| Policy readout recoverability What a linear readout of the policy's own residual stream recovers, against the black-box bank on the same items. | What access it needs: forward, policy What kind of object: neural_gen, neural_scalar When it can run: post_run one forward pass per item plus a ridge solve; no grader calls Grader white-box: the existing battery, re-typed | implemented |
| Decision coverage What fraction of the verifier's logic any rollout has ever exercised. | What access it needs: grader, record What kind of object: program When it can run: post_run, pre_run Verifier and environment science | implemented |
| Surviving mutants Would the verifier notice if it were wrong? | What access it needs: grader, mutate Verifier and environment science | implemented |
| Metamorphic violations Relations the grader should respect and does not. | What access it needs: grader, query Verifier and environment science | implemented |
| Sensitivity indices Which inputs actually move the score. | What access it needs: grader, query Verifier and environment science | implemented |
| False-positive fuzzing Where the verifier accepts a wrong answer. | What access it needs: grader, query Verifier and environment science | implemented |
| Exploit-family coverage How much of the exploit space you have found. | What access it needs: not stated Verifier and environment science | implemented |
| The grader card The composite artifact and the wedge product. | What access it needs: grader, query, replicate Verifier and environment science | implemented |
| Attack surface What the harness exposes. | What access it needs: not stated Verifier and environment science | implemented |
| Static structure The verifier's shape, extracted. | What access it needs: not stated Verifier and environment science | implemented |
| Determinism and replay fidelity | What access it needs: query, task Verifier and environment science | implemented |
| Estimator specification, recorded What transform actually ran. | What access it needs: record The estimator: where a good reward becomes a bad gradient | implemented |
| Degenerate and all-fail group fraction | What access it needs: record The estimator: where a good reward becomes a bad gradient | implemented |
| Signal-versus-noise share of the gradient, and its attribution | What access it needs: record The estimator: where a good reward becomes a bad gradient | implemented |
| Amplifier safety Is this reward component safe to add? | What access it needs: record The estimator: where a good reward becomes a bad gradient | implemented |
| Clip and mask accounting | What access it needs: record The estimator: where a good reward becomes a bad gradient | implemented |
| Train-inference logprob mismatch | What access it needs: record The estimator: where a good reward becomes a bad gradient | implemented |
| The selection term and the residual | What access it needs: record The four books | implemented |
| The selection-explained fraction Λ The cheapest strong diagnostic here. | What access it needs: record The four books | implemented |
| The cost book: KL budget and efficiency | What access it needs: backward, policy, record The four books | implemented |
| The reconciliation residual, budgeted | What access it needs: not stated The four books | implemented |
| The selection covector and spectrum | What access it needs: backward, policy The four books | implemented |
| Realised versus predicted response, the Lande slope | What access it needs: not stated The four books | implemented |
| The credit measure | What access it needs: not stated Credit geometry | implemented |
| Turn and tool-call disintegration The agentic form. | What access it needs: record Credit geometry | implemented |
| The implicit process reward model, extracted | What access it needs: record Credit geometry | implemented |
| Localiser scoring, done validly | What access it needs: not stated Credit geometry | open |
| Successor representation | What access it needs: not stated Credit geometry | open |
| Adiabaticity number Is this run quasi-static? | What access it needs: record Rate and regime | implemented |
| The two-run rate test The decisive, cheap experiment. | What access it needs: control Rate and regime | implemented |
| Rate-extrapolated hysteresis | What access it needs: control Rate and regime | implemented |
| Relaxation time and the transition window | What access it needs: record Rate and regime | implemented |
| Regime reading The whole envelope, measured per step. | What access it needs: record Rate and regime | implemented |
| Density discontinuity at a reward gate McCrary as a hacking detector. | What access it needs: record Adversarial pressure, and the monitor as a target | implemented |
| Bunching elasticity How hard the policy is pushing on the gate. | What access it needs: record Adversarial pressure, and the monitor as a target | implemented |
| Gradient dead-zone fraction | What access it needs: record Adversarial pressure, and the monitor as a target | implemented |
| Monitor degradation under pressure The figure of merit nobody uses. | What access it needs: control Adversarial pressure, and the monitor as a target | implemented |
| The derivative of within-group reward variance | What access it needs: record Adversarial pressure, and the monitor as a target | implemented |
| Obfuscation pressure on the monitored channel | What access it needs: not stated Adversarial pressure, and the monitor as a target | open |
| Exploit transfer graph | What access it needs: not stated Adversarial pressure, and the monitor as a target | open |
| E-values and confidence sequences Peek freely. | What access it needs: record Monitoring | implemented |
| ARL-designed alarms | What access it needs: record Monitoring | implemented |
| Conjunction detectors | What access it needs: record Monitoring | implemented |
| Operating point from an asymmetric loss | What access it needs: not stated Monitoring | implemented |
| Check standard | What access it needs: not stated Monitoring | implemented |
| The forecast calibration record, scored against the library's own forecasts The score of the library's own forecasts, in public, with the honest starting position printed at the top. | What access it needs: record What kind of object: composite, human, neural_gen, neural_scalar, procedural, program When it can run: post_run none Monitoring | implemented |
| The distillation gap Kill risk number one. | What access it needs: not stated Transfer and survival | open |
| Planted-to-real transfer coefficient | What access it needs: not stated Transfer and survival | open |
| Shelf life of a readout | What access it needs: not stated Transfer and survival | open |
| Sparsity under staleness | What access it needs: not stated Transfer and survival | open |
| Reference-material certificate | What access it needs: mutate, organism Reference materials and label validity | implemented |
| Label error rate and the ceiling it implies | What access it needs: not stated Reference materials and label validity | implemented |
| Two-sided verifier error | What access it needs: not stated Reference materials and label validity | implemented |
| The distributed-tell question, which is a free white-box experiment | What access it needs: forward, policy Reference materials and label validity | implemented |
| Position-stratified null for localisation | What access it needs: verif Reference materials and label validity | implemented |
| Substrate noise floor, LOD and LOQ §4.7. Consulted by every preflight, cached per configuration. | What access it needs: not stated Meta-instruments: measurements about measurements | implemented |
| Instrument effect The overhead this measurement imposed, per step, as a term in the uncertainty budget. **No competitor has published an overhead number of any kind**, so a monitor with a measured one would be the first. | What access it needs: not stated Meta-instruments: measurements about measurements | implemented |
| Dumb-baseline bank String match, length, TF-IDF, n-gram diversity, a scaffolded black-box prompt, the gradient-norm peak. Every claim ships against all six. The reason is a single published case: a probe reported at AUC 0.998 on a task where a zero-parameter string match scores 100%, whose author's summary is "the probe detects the hack, and the detection is empty." | What access it needs: not stated Meta-instruments: measurements about measurements | implemented |
| Semantic placebo A coherent but irrelevant direction, not a random Gaussian one. Mandatory on every steering and ablation claim, because a vampires-versus-werewolves direction suppressed deployment-time hacking to 0.000 exactly as well as the real direction. | What access it needs: not stated Meta-instruments: measurements about measurements | implemented |
| Matched positive control A null without an identically-powered positive control is indistinguishable from an underpowered experiment. `NO_MATCHED_CONTROL` is a refusal reason. The library's own susceptibility card at accepted power 0.13 conspicuously lacked one. | What access it needs: not stated Meta-instruments: measurements about measurements | implemented |
| Reasoning-stripped ablation For anything that reads text, because a leading detector drops from AUC 0.9467 to 0.6213 once natural-language reasoning is removed from its input, and reviewers will run this if we do not. | What access it needs: not stated Meta-instruments: measurements about measurements | open |
| Uncertainty budget The GUM table, lint-enforced to compose, with the largest term named. It is almost never sampling noise. | What access it needs: not stated Meta-instruments: measurements about measurements | implemented |
| Interlaboratory comparison HuggingFace against vLLM, eager against compiled, two seeds, two precisions. Report `s_L` and the Birge ratio beside every reading. | What access it needs: not stated Meta-instruments: measurements about measurements | implemented |
| Incremental validity §6.4. Required on every white-box reading. | What access it needs: not stated Meta-instruments: measurements about measurements | implemented |
| Power and minimum detectable effect Before the run, not after. Three of five standard power calculators are roughly 2x wrong for close paired model comparisons, and the resolution ratio `q = N/N*` says outright when a leaderboard row is not resolved. Also carry the detection-band lesson: position bias is statistically detectable only within roughly a 60% to 95% base-accuracy window, so **absence of signal above that band should be read as "not measurable", not as "unbiased"**, and that logic applies to several of the library's own refuted cards. | What access it needs: not stated Meta-instruments: measurements about measurements | implemented |
| Estimator disagreement, published When two estimators of the same quantity disagree on the same data, that difference is the cheap method's transfer uncertainty. It becomes a `Transfer` row and it composes into the calibration chain, measured against the same corpus. Nobody publishes this. | What access it needs: not stated Meta-instruments: measurements about measurements | implemented |
| The reward-versus-gold frontier The whole frontier is estimable before any optimisation is run, from n base-policy rollouts scored by both channels. | What access it needs: gold, grader, policy, query What kind of object: composite, human, neural_gen, neural_scalar, procedural, program When it can run: pre_run n grader calls and n gold calls, both reusable across every lambda on the sweep The frontier: what would happen if we optimised | implemented |
| The visibility horizon Where a tilt extrapolation goes blind, which nobody states. | What access it needs: grader, policy, query What kind of object: composite, human, neural_gen, neural_scalar, procedural, program When it can run: pre_run The frontier: what would happen if we optimised | implemented |
| Reward tail index The precondition the whole tilt layer rests on, measured rather than assumed. | What access it needs: grader, policy, query What kind of object: composite, human, neural_gen, neural_scalar, procedural, program When it can run: pre_run The frontier: what would happen if we optimised | implemented |
| Surrogate conditions and the concomitant of best-of-n Falsifiable conditions on the proxy beat a predicted turning point, and best-of-n Goodhart is a concomitant problem nobody has recognised as one. | What access it needs: gold, grader, query What kind of object: composite, human, neural_gen, neural_scalar, procedural, program When it can run: pre_run The frontier: what would happen if we optimised | implemented |
| Optimal component weights Nobody sets a reward component's weight from that component's measured noise, and the formula that does has been sitting in contract theory since 1991. | What access it needs: grader, replicate What kind of object: composite, human, neural_gen, neural_scalar, procedural, program When it can run: post_run, pre_run The frontier: what would happen if we optimised | implemented |
| The equal-compensation table Which component of a composite reward the policy is being paid least to work on. | What access it needs: control, grader, policy, record What kind of object: composite, human, neural_gen, neural_scalar, procedural, program When it can run: post_run, pre_run free. Arithmetic on parameters the caller already holds. The frontier: what would happen if we optimised | implemented |
| The sorting cutoff Do not put a noisy judge and a crisp unit test in the same weighted sum, with the threshold at which that stops being a slogan. | What access it needs: control, grader, policy, replicate What kind of object: composite, human, neural_gen, neural_scalar, procedural, program When it can run: post_run, pre_run free. Arithmetic on parameters the caller already holds. The frontier: what would happen if we optimised | implemented |
| Noise and angle, per component Every reward component needs at least two numbers, a noise and an angle, and no RLHF tooling separates them. | What access it needs: control, grader, policy, replicate What kind of object: composite, human, neural_gen, neural_scalar, procedural, program When it can run: post_run, pre_run free. Arithmetic on parameters the caller already holds. The frontier: what would happen if we optimised | implemented |
What is not here
The registry defines 190 quantities. 43 have an estimator. The remaining 147 are open research targets, and they are the same list that drives the open problems page.
A catalogue with no stated gap reads as complete. Publishing the open count is what makes the implemented count mean anything.
9 of the names in the table above are not the ones the library's registry uses. Each is printed below with the field it replaced, because a name changed silently is a name a reader cannot check against the package. The registry's own names ship inside the library, in its instrument catalogue.
- Signal validity: the grader as a measurement device
- series title
- Reference materials and label validity
- series title
- The forecast calibration record, scored against the library's own forecasts
- instrument name
- Estimator disagreement, published
- instrument name
- Reference materials and calibration
- area name
- Attribution versus patching, resolved against planted ground truth
- card title
- Instrument calibration transfers from a CPU organism to a real 0.6B trunk
- card title
- The v2 campaign scores its own calibration against a coin
- card title
- When two estimators of the same quantity disagree on the same data, that difference is the cheap method's transfer uncertainty. It becomes a `Transfer` row and it composes into the calibration chain, measured against the same corpus. Nobody publishes this.
- instrument headline