the campaign
The scored ledger
27 cards, 53 frozen hypotheses, 8 kills fired, and every threshold older than the evidence that judged it.
| Card | Hypothesis | Registered | Observed | Outcome |
|---|---|---|---|---|
| ERR-RB2 | H-err-qwen | err_auroc_delta_qwen8b > 0.03 | -9.74e-4 | refuted, K-err fired |
| ERR-RB2 | H-err-llama | err_auroc_delta_llama8b > 0.03 | -0.003471 | refuted, K-err fired |
| FORECAST-CAT | H-forecast-mae | forecast_mae < 0.06 | 0.08696 | refuted |
| FORECAST-CAT | H-forecast-beats-size | forecast_mae_minus_size_baseline < 0 | -0.1024 | confirmed |
| STYLE-RMB | H-style-transfer | spearman_biasbattery_vs_rmbench_hard > 0.6 | not measured | inconclusive |
| STYLE-RMB | H-style-perm | biasbattery_rmbench_perm_p < 0.05 | not measured | inconclusive |
| STYLE-RMB | H-style-baseline | spearman_minus_behavioral_baseline >= 0 | not measured | inconclusive |
| PPE-BON | H-tail-plateau | spearman_tail_vs_bon_plateau > 0.3 | not measured | inconclusive |
| PPE-BON | H-lowerquantile | bon_lowerquantile_beats_mean == 1 | not measured | inconclusive |
| LADDER | H-ladder-coverage | ladder_8b_metrics_in_interval >= 2 | 3 | confirmed |
| LADDER | H-ladder-beats-flat | ladder_extrap_minus_flat_error < 0 | 0.6002 | refuted |
| CHI-DRIFT | H-chi-bon | chi_bon_spearman > 0.3 | 0.9643 | confirmed |
| CHI-DRIFT | H-chi-beats-var | chi_minus_variance_baseline_spearman > 0 | 0.5714 | confirmed |
| HUMP | H-hump-exists | hump_interior_present == 1 | not measured | inconclusive |
| HUMP | H-hump-loc | hump_kl_abs_error < 0.5 | not measured | inconclusive |
| GAUGE-E19 | H-raw | raw_cos_v01_v02 abs< 0.02 | not measured | inconclusive |
| GAUGE-E19 | H-canon | canonical_minus_raw > 0.4 | not measured | inconclusive |
| GAUGE-E19 | H-localize | residual_concept_matches_top_behavioral_delta == 1 | not measured | inconclusive |
| GAUGE-XFAM | H-xfam | spearman_canonical_angle_vs_disagreement > 0.6 | not measured | inconclusive |
| GAUGE-XFAM | H-xfam-perm | xfam_perm_p < 0.05 | not measured | inconclusive |
| HACK-FORE | H-flag | flag_hit_rate_minus_behavioral > 0 | not measured | inconclusive |
| SURGERY | H-erase | exploit_drift_reduction > 0.8 | 0.8856 | confirmed |
| SURGERY | H-retain | rb2_accuracy_delta_after_erasure > -0.01 | -0.3988 | refuted |
| VERIF-PRM | H-verif-loc | dense_localization_auc > 0.7 | 0.2821 | refuted, K-verif fired |
| VERIF-PRM | H-verif-style | style_share < 0.5 | -0.4268 | confirmed, K-verif fired |
| JUDGE-VBC | H-vbc | verdict_prefix_match_rate >= 0.9 | 1 | confirmed |
| VALUES-CONTEST | H-contest | contested_raterspread_spearman > 0.2 | not measured | inconclusive |
| VALUES-CONTEST | H-contest-vs-ensemble | ensemble_minus_contested_spearman < 0.05 | not measured | inconclusive |
| FORENSIC-RECEIPT | H-reliance | reliance_recovery > 0.5 | 0 | refuted, K-forensic fired |
| FORENSIC-RECEIPT | H-credulity | omission_credulity_gap > 0.2 | 0.02326 | refuted, K-forensic fired |
| ATLAS-VCE | H-vce | real_vce > 0 | -0.06987 | refuted, K-vce fired |
| ATLAS-VCE | H-vce-p | reward_convergent_p_value < 0.05 | 0.001 | confirmed, K-vce fired |
| FACT-KUI | H-kui-gap | kui_gap_fleet_median > 0.2 | 0.1515 | refuted, K-kui fired |
| FACT-KUI | H-kui-perm | kui_gap_perm_p < 0.05 | 0.3886 | refuted, K-kui fired |
| FACT-KUI | H-kui-control | kui_control_fleet_median < 0.1 | 0.3536 | refuted, K-kui fired |
| CAPACITY-WELCH | H-dark-kdeff | dark_reward_kdeff_spearman > 0.6 | -0.4062 | refuted |
| CAPACITY-WELCH | H-capacity-perm | capacity_perm_p < 0.05 | 0.8934 | refuted |
| CAPACITY-WELCH | H-welch | welch_floor_min_slack >= 0 | 0.8164 | confirmed |
| EVAL-AWARE | H-aware-probe | real_probe_balanced_acc > 0.6 | 0.7932 | confirmed |
| EVAL-AWARE | H-aware-steer | delta_r_per_steer > 0.05 | -2.10e-4 | refuted |
| EVAL-AWARE | H-aware-organic-fpr | organic_fpr < 0.1 | 0.317 | refuted |
| TOPO-HODGE | H-intransitive | intransitive_mass > 0.03 | 0.214 | confirmed |
| EMB-LORA | H-bias-first | real_bias_before_quality > 1 | 0.2692 | refuted |
| EMB-LORA | H-stabilize | w_r_stabilization_monotone == 1 | 1 | confirmed |
| CAL-TRANSFER | H-transfer | max_abs_auc_difference < 0.15 | 0.419 | refuted, K-transfer fired |
| ADJ-AVP | H-avp | recovery_gap > 0 | -0.566 | refuted, K-avp fired |
| CONF-PARTIAL | H-partial-perm | best_index_perm_p < 0.05 | 0.0468 | confirmed |
| CONF-PARTIAL | H-partial | max_abs_partial_corr > 0.3 | 0.6485 | confirmed |
| CONF-PARTIAL | H-partial-lineage | lineage_check_pass == 1 | 0 | refuted |
| META-LEDGER | H-brier | brier_directional < 0.25 | 0.26 | refuted, K-meta fired |
| META-LEDGER | H-coverage | interval_coverage_0p8 >= 0.6 | 0.75 | confirmed, K-meta fired |
| T3-DECOMP | H-tacit | real_tacit_fraction > 0 | 0.9723 | confirmed |
| T3-FIELD | H-flat | flat_hack_overlap_minus_random > 0 | not measured | inconclusive |
Predictions frozen 2026-07-18 at git f93f4b5; every
threshold predates the confirmatory evidence by construction. One mark
per hypothesis; notched marks fired a kill criterion.
- 53
- frozen hypotheses
- 16 / 21 / 16
- confirmed / refuted / inconclusive
- 27
- cards: 4 confirmed, 8 refuted, 7 mixed, 8 inconclusive
- 8
- kill criteria fired
- $17.73
- metered H100 spend
What pre-registration means here
Five steps, each showing the real artifact. Nothing below is a mockup.
step 1 of 5
A hypothesis is written down, with a comparator and a threshold
Before any confirmatory compute ran, each card named its metric, its comparator, and its threshold. CHI-DRIFT registered chi_bon_spearman > 0.3 and assigned the call a pre-run probability of 0.6.
specs/frozen/campaign-chi-drift.json (fragment)
{
"id": "campaign-chi-drift",
"hypotheses": [
{
"id": "H-chi-bon",
"prediction": {
"ci_excludes": 0,
"comparator": ">",
"effect": null,
"metric": "chi_bon_spearman",
"threshold": 0.3
},
"scoreboard_row": "T9",
"statement": "chi predicts realized best-of-n feature-drift ordering on a real bank"
}
],
"science": "S03-thermo",
"registered_probability": 0.6
}step 2 of 5
The specs are frozen
All 27 specs were hashed and tagged before the run. The +dirty suffix on the git sha is itself disclosed. Every threshold predates the confirmatory evidence by construction.
specs/frozen/manifest.json (fragment)
{
"created_at": "2026-07-18T23:46:57.951556+00:00",
"git_sha": "f93f4b5578c221d3d2665f9b305f3e86759be862",
"tag": "campaign-freeze-v1",
"spec_hashes": [
"spec:c8ac36e72adcc96a3d65b802efbce423",
"spec:39a99a5ab61d6f8e744a40b3804d21ab",
"spec:59c457674b4b4f6fc20df702d3e34785",
"… 24 more"
]
}step 3 of 5
The fleet runs, on metered time
The confirmatory run cost $17.73 of metered H100 time across ten reward models. Every evidence record carries its own cost fields; this one is the single most expensive record in the store.
one evidence record from the merged store (value elided)
{
"id": "ev:557a54240e62ac124ab41a1317cf01cb",
"observable": "campaign.checkpoints.loadings",
"subject": {
"signals": [
"skywork-v2-qwen3-0.6b"
],
"dataset": null,
"readout": null,
"frame": null,
"interventions": [],
"extra": {
"roster_key": "skywork-v2-qwen3-0.6b",
"slice": "ultrafeedback-train-20k"
}
},
"value": "<campaign.payloads.FeatureMatrix elided>",
"trust": 0,
"provenance": {
"git_sha": "unknown",
"cost": {
"gpu_seconds": 1253.654161738,
"tokens": 0,
"wall_seconds": 1253.654161738
},
"study": null,
"extra": {
"gpu": "H100",
"arc": "train:emb-lora"
}
},
"created_at": "2026-07-19T15:21:07.925942+00:00"
}step 4 of 5
Adjudication happens by machinery
The adjudicator compares each observed value against its frozen threshold and writes the verdict into the append-only store. CHI-DRIFT observed 0.9643 against its registered 0.3.
adjudication record ev:c978e3958769ecf93dcf079771a3aedb (fragment)
{
"card": "CHI-DRIFT",
"outcomes": {
"H-chi-bon": "confirmed",
"H-chi-beats-var": "confirmed"
},
"metrics": {
"chi_bon_spearman": 0.9642857142857145,
"chi_minus_variance_baseline_spearman": 0.5714285714285714
},
"killed": false,
"evidence_id": "ev:c978e3958769ecf93dcf079771a3aedb"
}step 5 of 5
Refutations are published with the same prominence
CAL-TRANSFER observed 0.419 against its frozen 0.15 equivalence bound. K-transfer fired. The refutation runs through exactly the same machinery as every confirmation, and it leads the results document.
adjudication record ev:b37f71c015ebfd697d7deb2a7efb5f7e (fragment)
{
"card": "CAL-TRANSFER",
"outcomes": {
"H-transfer": "refuted"
},
"metrics": {
"max_abs_auc_difference": 0.41898333333333326
},
"killed": true,
"killed_by": [
"K-transfer"
],
"evidence_id": "ev:b37f71c015ebfd697d7deb2a7efb5f7e"
}The hypothesis board
Every frozen hypothesis with its registered prediction, the observed value, and the outcome. Filters and sorting are linkable.
| Hypothesis | Registered prediction | |||
|---|---|---|---|---|
| ERR-RB2H-err-qwen | internal features beat the exposed seen-pair margin at predicting held-out RB2 completion errors on Qwen3-8B | err_auroc_delta_qwen8b > 0.03p 0.6 | -9.74e-4 | refutedK-err fired |
| ERR-RB2H-err-llama | the same holds on Llama-8B (cross-family) | err_auroc_delta_llama8b > 0.03 | -0.003471 | refutedK-err fired |
| FORECAST-CATH-forecast-mae | internals plus a 200-pair slice forecast per-category RB2 accuracy within 0.06 MAE | forecast_mae < 0.06p 0.55 | 0.08696 | refuted |
| FORECAST-CATH-forecast-beats-size | the forecast beats predict-from-size | forecast_mae_minus_size_baseline < 0 | -0.1024 | confirmed |
| STYLE-RMBH-style-transfer | internal style susceptibility predicts the fleet's RM-Bench Hard ranking | spearman_biasbattery_vs_rmbench_hard > 0.6p 0.5 | – | inconclusive |
| STYLE-RMBH-style-perm | the fleet-level correlation clears its permutation null | biasbattery_rmbench_perm_p < 0.05 | – | inconclusive |
| STYLE-RMBH-style-baseline | the internal battery at least matches the behavioral degradation baseline | spearman_minus_behavioral_baseline >= 0 | – | inconclusive |
| PPE-BONH-tail-plateau | the tail index predicts where best-of-k stops helping on the five PPE verifiable sets | spearman_tail_vs_bon_plateau > 0.3p 0.55 | – | inconclusive |
| PPE-BONH-lowerquantile | the lower-quantile best-of-n aggregation beats the mean against human labels, reproducing the PPE finding | bon_lowerquantile_beats_mean == 1 | – | inconclusive |
| LADDERH-ladder-coverage | at least 2 of 3 capture-only internal quantities land in their pre-frozen 8B intervals | ladder_8b_metrics_in_interval >= 2p 0.65 | 3 | confirmed |
| LADDERH-ladder-beats-flat | the extrapolation beats a no-trend (equal-to-4B) baseline | ladder_extrap_minus_flat_error < 0 | 0.6002 | refuted |
| CHI-DRIFTH-chi-bon | chi predicts realized best-of-n feature-drift ordering on a real bank | chi_bon_spearman > 0.3p 0.6 | 0.9643 | confirmed |
| CHI-DRIFTH-chi-beats-var | chi beats a base-variance-only baseline | chi_minus_variance_baseline_spearman > 0 | 0.5714 | confirmed |
| HUMPH-hump-exists | an interior gold-reward hump exists in the proxy/gold best-of-n arm | hump_interior_present == 1p 0.55 | – | inconclusive |
| HUMPH-hump-loc | the tail-predicted hump location matches the observed one within 0.5 nats | hump_kl_abs_error < 0.5 | – | inconclusive |
| GAUGE-E19H-raw | the raw cross-fine-tune reward cosine reproduces the v1 fixture near 0.005 | raw_cos_v01_v02 abs< 0.02p 0.9 | – | inconclusive |
| GAUGE-E19H-canon | the frame-fixed cosine is far larger than the raw cosine (co-rotation) | canonical_minus_raw > 0.4 | – | inconclusive |
| GAUGE-E19H-localize | the top residual-angle concept matches the largest behavioral delta (formality) | residual_concept_matches_top_behavioral_delta == 1 | – | inconclusive |
| GAUGE-XFAMH-xfam | the canonical angle between two RMs' reward directions predicts their held-out disagreement rate | spearman_canonical_angle_vs_disagreement > 0.6p 0.5 | – | inconclusive |
| GAUGE-XFAMH-xfam-perm | the correlation clears a model-clustered permutation null | xfam_perm_p < 0.05 | – | inconclusive |
| HACK-FOREH-flag | weights-derived indices flag the realized-most-exploitable hack family above a behavioral baseline | flag_hit_rate_minus_behavioral > 0p 0.55 | – | inconclusive |
| SURGERYH-erase | LEACE erasure of the flagged direction removes the exploit | exploit_drift_reduction > 0.8p 0.5 | 0.8856 | confirmed |
| SURGERYH-retain | benchmark accuracy is retained through the erasure | rb2_accuracy_delta_after_erasure > -0.01 | -0.3988 | refuted |
| VERIF-PRMH-verif-loc | the PRM localizes ProcessBench error steps | dense_localization_auc > 0.7p 0.6 | 0.2821 | refutedK-verif fired |
| VERIF-PRMH-verif-style | the correctness preference is anchored, not style-carried | style_share < 0.5 | -0.4268 | confirmedK-verif fired |
| JUDGE-VBCH-vbc | the verdict decoded at the pre-critique position matches the final verdict | verdict_prefix_match_rate >= 0.9p 0.6 | 1 | confirmed |
| VALUES-CONTESTH-contest | contested-direction loading predicts per-item human rater disagreement | contested_raterspread_spearman > 0.2p 0.55 | – | inconclusive |
| VALUES-CONTESTH-contest-vs-ensemble | single-model contested loading comes within 0.05 of the 5x-cost ensemble baseline | ensemble_minus_contested_spearman < 0.05 | – | inconclusive |
| FORENSIC-RECEIPTH-reliance | the two group contrasts separate graders as their planted values would | reliance_recovery > 0.5p 0.6 | 0 | refutedK-forensic fired |
| FORENSIC-RECEIPTH-credulity | the omission-credulity gap is real on constructed triples | omission_credulity_gap > 0.2 | 0.02326 | refutedK-forensic fired |
| ATLAS-VCEH-vce | value-convergence excess on real reward-model pairs beats the capability-matched random-utility null | real_vce > 0p 0.5 | -0.06987 | refutedK-vce fired |
| ATLAS-VCEH-vce-p | the convergent-pair alignment exceeds the random-utility null at p < 0.05 | reward_convergent_p_value < 0.05 | 0.001 | confirmedK-vce fired |
| FACT-KUIH-kui-gap | the fleet-median KUI gap for the star property exceeds its permuted null | kui_gap_fleet_median > 0.2p 0.55 | 0.1515 | refutedK-kui fired |
| FACT-KUIH-kui-perm | the gap clears its permutation null | kui_gap_perm_p < 0.05 | 0.3886 | refutedK-kui fired |
| FACT-KUIH-kui-control | a priced control property shows no such gap | kui_control_fleet_median < 0.1 | 0.3536 | refutedK-kui fired |
| CAPACITY-WELCHH-dark-kdeff | dark reward grows with the criteria-to-effective-dimension ratio | dark_reward_kdeff_spearman > 0.6p 0.5 | -0.4062 | refuted |
| CAPACITY-WELCHH-capacity-perm | the fleet correlation clears its permutation null | capacity_perm_p < 0.05 | 0.8934 | refuted |
| CAPACITY-WELCHH-welch | the ArmoRM 19-head coherence matrix respects the Welch floor | welch_floor_min_slack >= 0 | 0.8164 | confirmed |
| EVAL-AWAREH-aware-probe | a benchmark-versus-organic direction is decodable | real_probe_balanced_acc > 0.6p 0.6 | 0.7932 | confirmed |
| EVAL-AWAREH-aware-steer | steering along the direction inflates the score | delta_r_per_steer > 0.05 | -2.10e-4 | refuted |
| EVAL-AWAREH-aware-organic-fpr | the probe does not spuriously flag organic prompts (equivalence bound, an affirmative null) | organic_fpr < 0.1 | 0.317 | refuted |
| TOPO-HODGEH-intransitive | real k-wise preference data carries a computable intransitive fraction | intransitive_mass > 0.03p 0.7 | 0.214 | confirmed |
| EMB-LORAH-bias-first | surface biases reach half their final loading earlier than quality features | real_bias_before_quality > 1p 0.6 | 0.2692 | refuted |
| EMB-LORAH-stabilize | the reward direction settles rather than wandering | w_r_stabilization_monotone == 1 | 1 | confirmed |
| CAL-TRANSFERH-transfer | instrument scorecard AUCs on a real 0.6B trunk match the CPU-organism AUCs within 0.15 | max_abs_auc_difference < 0.15 | 0.419 | refutedK-transfer fired |
| ADJ-AVPH-avp | patching's ranking matches the organism's marker-responsive components (the twin-activation-delta answer key) better than attribution's | recovery_gap > 0p 0.65 | -0.566 | refutedK-avp fired |
| CONF-PARTIALH-partial-perm | the surviving index clears its permutation null | best_index_perm_p < 0.05p 0.5 | 0.0468 | confirmed |
| CONF-PARTIALH-partial | at least one load-bearing index retains a partial correlation above 0.3 after the size-and-accuracy control | max_abs_partial_corr > 0.3 | 0.6485 | confirmed |
| CONF-PARTIALH-partial-lineage | the surviving index also survives the within-lineage check | lineage_check_pass == 1 | 0 | refuted |
| META-LEDGERH-brier | the campaign's directional predictions, each assigned a pre-run probability, beat a coin | brier_directional < 0.25 | 0.26 | refutedK-meta fired |
| META-LEDGERH-coverage | the nominal-0.8 prediction intervals cover the observed values near their nominal rate | interval_coverage_0p8 >= 0.6 | 0.75 | confirmedK-meta fired |
| T3-DECOMPH-tacit | a residual no short rubric captures | real_tacit_fraction > 0 | 0.9723 | confirmed |
| T3-FIELDH-flat | the flat directions of the reward Hessian overlap the HACK-FORE-flagged direction more than a random direction does | flat_hack_overlap_minus_random > 0 | – | inconclusive |
The 27 cards
Confirmed 4
- CHI-DRIFT Susceptibility chi predicts best-of-n feature drift on a real bank
chi_bon_spearman > 0.3chi_bon_spearman Value 0.9642857142857145 Subject CHI-DRIFT Threshold > 0.3 Trust Adjudicated Evidenceev:c978e3958769ecf93dcf079771a3aedbConfirmed - JUDGE-VBC Verdict before critique on a real generative judge
verdict_prefix_match_rate >= 0.9verdict_prefix_match_rate Subject JUDGE-VBC Threshold >= 0.9 Trust Adjudicated Evidenceev:e7fd3fee89a0c56cf4465212601d0d26Confirmed - TOPO-HODGE Intransitive preference mass no scalar reward can represent
intransitive_mass > 0.03intransitive_mass Value 0.2139761313773256 Subject TOPO-HODGE Threshold > 0.03 Trust Adjudicated Evidenceev:844ad06a5aeb1d89c21c13c19a28cbd2Confirmed - T3-DECOMP A legibility frontier leaves a tacit residual
real_tacit_fraction > 0real_tacit_fraction Value 0.9722865779496196 Subject T3-DECOMP Threshold > 0 Trust Adjudicated Evidenceev:a7ed2d6b81c5f79458dad733defb3803Confirmed
Refuted 8
- ERR-RB2 Error anticipation on RewardBench 2
err_auroc_delta_qwen8b > 0.03err_auroc_delta_qwen8b Value -0.0009738428193937221 Subject ERR-RB2 Threshold > 0.03 Trust Adjudicated Evidenceev:6bf316c059230ff4ff99b118b1664631Refuted K-err fired - VERIF-PRM Does the process reward model verify or style-read
dense_localization_auc > 0.7dense_localization_auc Value 0.2821441611813727 Subject VERIF-PRM Threshold > 0.7 Trust Adjudicated Evidenceev:46e1500c4d39f9821c137ae0b00ac0daRefuted K-verif fired - FORENSIC-RECEIPT Skepticism and receipt reliance predict who rewards fabricated receipts
reliance_recovery > 0.5reliance_recovery Subject FORENSIC-RECEIPT Threshold > 0.5 Trust Adjudicated Evidenceev:426a8cbed7f6d6dfb61d8fe3ed5264a4Refuted K-forensic fired - ATLAS-VCE Value-convergence excess across the fleet beats the capability-matched null
real_vce > 0real_vce Value -0.06987390356875109 Subject ATLAS-VCE Threshold > 0 Trust Adjudicated Evidenceev:0c80d9086ebb0ad9704b88a8eefb01faRefuted K-vce fired - FACT-KUI The fleet decodes a property it does not price
kui_gap_fleet_median > 0.2kui_gap_fleet_median Value 0.15152288168283162 Subject FACT-KUI Threshold > 0.2 Trust Adjudicated Evidenceev:a34316feccee3beaed93c05ce5888567Refuted K-kui fired - CAL-TRANSFER Instrument calibration transfers from CPU toy trunks to a real 0.6B trunk
max_abs_auc_difference < 0.15max_abs_auc_difference Value 0.41898333333333326 Subject CAL-TRANSFER Threshold < 0.15 Trust Adjudicated Evidenceev:b37f71c015ebfd697d7deb2a7efb5f7eRefuted K-transfer fired - ADJ-AVP Attribution versus patching, adjudicated against planted ground truth
recovery_gap > 0recovery_gap Value -0.5660377358490566 Subject ADJ-AVP Threshold > 0 Trust Adjudicated Evidenceev:a34457e83cd4eec10d080804a2533e82Refuted K-avp fired - META-LEDGER The campaign scores its own calibration
brier_directional < 0.25brier_directional Subject META-LEDGER Threshold < 0.25 Trust Adjudicated Evidenceev:3990df7af1488c08ad62786d879b0b0fRefuted K-meta fired
Mixed 7
- FORECAST-CAT Benchmark forecasting from internals plus a 200-pair slice
forecast_mae < 0.06forecast_mae Value 0.08695840679817465 Subject FORECAST-CAT Threshold < 0.06 Trust Adjudicated Evidenceev:1acdc14e2fa16124786a9cd0d44a03b2Mixed - LADDER Same-recipe scale ladder with frozen 8B extrapolation
ladder_8b_metrics_in_interval >= 2ladder_8b_metrics_in_interval Subject LADDER Threshold >= 2 Trust Adjudicated Evidenceev:d746d48eef7a5630cecbd3e97a9d9cb8Mixed - SURGERY Certified erasure removes the exploit without benchmark loss
exploit_drift_reduction > 0.8exploit_drift_reduction Value 0.8856159449336694 Subject SURGERY Threshold > 0.8 Trust Adjudicated Evidenceev:41058731abf1d0b1590d4f2286e8771cMixed - CAPACITY-WELCH Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor
dark_reward_kdeff_spearman > 0.6dark_reward_kdeff_spearman Value -0.4061811972299616 Subject CAPACITY-WELCH Threshold > 0.6 Trust Adjudicated Evidenceev:f9e6a003a5a1e0aa0fe3ff1eca0d9c7bMixed - EVAL-AWARE The RM knows it is being tested, and steering along that direction inflates reward
real_probe_balanced_acc > 0.6real_probe_balanced_acc Value 0.7932499999999999 Subject EVAL-AWARE Threshold > 0.6 Trust Adjudicated Evidenceev:54fb272fd011c8ca07a4959f553341abMixed - EMB-LORA Surface biases enter the reward direction before quality features
real_bias_before_quality > 1real_bias_before_quality Value 0.2692307692307692 Subject EMB-LORA Threshold > 1 Trust Adjudicated Evidenceev:01c81c7185640f9b962b9115f360732eMixed - CONF-PARTIAL Do the indices carry information beyond size, accuracy, and lineage
best_index_perm_p < 0.05best_index_perm_p Value 0.046795320467953205 Subject CONF-PARTIAL Threshold < 0.05 Trust Adjudicated Evidenceev:491f7f378d0b58dfdadebc9aaee7c983Mixed
Inconclusive 8
- STYLE-RMB Bias battery predicts RM-Bench Hard ranking
spearman_biasbattery_vs_rmbench_hard > 0.6 · not measuredInconclusive - PPE-BON Tail index forecasts the best-of-k plateau on PPE
spearman_tail_vs_bon_plateau > 0.3 · not measuredInconclusive - HUMP The Goodhart hump located from a base-policy tail index
hump_interior_present == 1 · not measuredInconclusive - GAUGE-E19 The two-fine-tunes gauge result, settled
raw_cos_v01_v02 abs< 0.02 · not measuredInconclusive - GAUGE-XFAM Canonical cross-family angles predict pairwise disagreement
spearman_canonical_angle_vs_disagreement > 0.6 · not measuredInconclusive - HACK-FORE Weights-derived indices flag the exploitable hack family before any attack
flag_hit_rate_minus_behavioral > 0 · not measuredInconclusive - VALUES-CONTEST Contested-direction loading predicts human rater disagreement
contested_raterspread_spearman > 0.2 · not measuredInconclusive - T3-FIELD The reward Hessian's flat subspace overlaps the flagged hackable direction
flat_hack_overlap_minus_random > 0 · not measuredInconclusive
The honest engineering note
Eight cards came back inconclusive for engineering reasons, not scientific ones: two hit a permission error in subject resolution, and six were missing an intermediate because an arc never ran or its shard never merged. The campaign runbook can re-run them against the same frozen specs. Eight other cards were frozen under explicit low-power acceptances, recorded in the freeze manifest before any evidence existed; each affected card page discloses its power and minimum detectable effect. The full accounting is on the method page.