The 53 adjudicated hypotheses: card, hypothesis, registered prediction, observed value, outcome.
CardHypothesisRegisteredObservedOutcome
ERR-RB2H-err-qwenerr_auroc_delta_qwen8b > 0.03-9.74e-4refuted, K-err fired
ERR-RB2H-err-llamaerr_auroc_delta_llama8b > 0.03-0.003471refuted, K-err fired
FORECAST-CATH-forecast-maeforecast_mae < 0.060.08696refuted
FORECAST-CATH-forecast-beats-sizeforecast_mae_minus_size_baseline < 0-0.1024confirmed
STYLE-RMBH-style-transferspearman_biasbattery_vs_rmbench_hard > 0.6not measuredinconclusive
STYLE-RMBH-style-permbiasbattery_rmbench_perm_p < 0.05not measuredinconclusive
STYLE-RMBH-style-baselinespearman_minus_behavioral_baseline >= 0not measuredinconclusive
PPE-BONH-tail-plateauspearman_tail_vs_bon_plateau > 0.3not measuredinconclusive
PPE-BONH-lowerquantilebon_lowerquantile_beats_mean == 1not measuredinconclusive
LADDERH-ladder-coverageladder_8b_metrics_in_interval >= 23confirmed
LADDERH-ladder-beats-flatladder_extrap_minus_flat_error < 00.6002refuted
CHI-DRIFTH-chi-bonchi_bon_spearman > 0.30.9643confirmed
CHI-DRIFTH-chi-beats-varchi_minus_variance_baseline_spearman > 00.5714confirmed
HUMPH-hump-existshump_interior_present == 1not measuredinconclusive
HUMPH-hump-lochump_kl_abs_error < 0.5not measuredinconclusive
GAUGE-E19H-rawraw_cos_v01_v02 abs< 0.02not measuredinconclusive
GAUGE-E19H-canoncanonical_minus_raw > 0.4not measuredinconclusive
GAUGE-E19H-localizeresidual_concept_matches_top_behavioral_delta == 1not measuredinconclusive
GAUGE-XFAMH-xfamspearman_canonical_angle_vs_disagreement > 0.6not measuredinconclusive
GAUGE-XFAMH-xfam-permxfam_perm_p < 0.05not measuredinconclusive
HACK-FOREH-flagflag_hit_rate_minus_behavioral > 0not measuredinconclusive
SURGERYH-eraseexploit_drift_reduction > 0.80.8856confirmed
SURGERYH-retainrb2_accuracy_delta_after_erasure > -0.01-0.3988refuted
VERIF-PRMH-verif-locdense_localization_auc > 0.70.2821refuted, K-verif fired
VERIF-PRMH-verif-stylestyle_share < 0.5-0.4268confirmed, K-verif fired
JUDGE-VBCH-vbcverdict_prefix_match_rate >= 0.91confirmed
VALUES-CONTESTH-contestcontested_raterspread_spearman > 0.2not measuredinconclusive
VALUES-CONTESTH-contest-vs-ensembleensemble_minus_contested_spearman < 0.05not measuredinconclusive
FORENSIC-RECEIPTH-reliancereliance_recovery > 0.50refuted, K-forensic fired
FORENSIC-RECEIPTH-credulityomission_credulity_gap > 0.20.02326refuted, K-forensic fired
ATLAS-VCEH-vcereal_vce > 0-0.06987refuted, K-vce fired
ATLAS-VCEH-vce-preward_convergent_p_value < 0.050.001confirmed, K-vce fired
FACT-KUIH-kui-gapkui_gap_fleet_median > 0.20.1515refuted, K-kui fired
FACT-KUIH-kui-permkui_gap_perm_p < 0.050.3886refuted, K-kui fired
FACT-KUIH-kui-controlkui_control_fleet_median < 0.10.3536refuted, K-kui fired
CAPACITY-WELCHH-dark-kdeffdark_reward_kdeff_spearman > 0.6-0.4062refuted
CAPACITY-WELCHH-capacity-permcapacity_perm_p < 0.050.8934refuted
CAPACITY-WELCHH-welchwelch_floor_min_slack >= 00.8164confirmed
EVAL-AWAREH-aware-probereal_probe_balanced_acc > 0.60.7932confirmed
EVAL-AWAREH-aware-steerdelta_r_per_steer > 0.05-2.10e-4refuted
EVAL-AWAREH-aware-organic-fprorganic_fpr < 0.10.317refuted
TOPO-HODGEH-intransitiveintransitive_mass > 0.030.214confirmed
EMB-LORAH-bias-firstreal_bias_before_quality > 10.2692refuted
EMB-LORAH-stabilizew_r_stabilization_monotone == 11confirmed
CAL-TRANSFERH-transfermax_abs_auc_difference < 0.150.419refuted, K-transfer fired
ADJ-AVPH-avprecovery_gap > 0-0.566refuted, K-avp fired
CONF-PARTIALH-partial-permbest_index_perm_p < 0.050.0468confirmed
CONF-PARTIALH-partialmax_abs_partial_corr > 0.30.6485confirmed
CONF-PARTIALH-partial-lineagelineage_check_pass == 10refuted
META-LEDGERH-brierbrier_directional < 0.250.26refuted, K-meta fired
META-LEDGERH-coverageinterval_coverage_0p8 >= 0.60.75confirmed, K-meta fired
T3-DECOMPH-tacitreal_tacit_fraction > 00.9723confirmed
T3-FIELDH-flatflat_hack_overlap_minus_random > 0not measuredinconclusive

Predictions frozen 2026-07-18 at git f93f4b5; every threshold predates the confirmatory evidence by construction. One mark per hypothesis; notched marks fired a kill criterion.

53
frozen hypotheses
16 / 21 / 16
confirmed / refuted / inconclusive
27
cards: 4 confirmed, 8 refuted, 7 mixed, 8 inconclusive
8
kill criteria fired
$17.73
metered H100 spend

What pre-registration means here

Five steps, each showing the real artifact. Nothing below is a mockup.

step 1 of 5

A hypothesis is written down, with a comparator and a threshold

Before any confirmatory compute ran, each card named its metric, its comparator, and its threshold. CHI-DRIFT registered chi_bon_spearman > 0.3 and assigned the call a pre-run probability of 0.6.

specs/frozen/campaign-chi-drift.json (fragment)

{
 "id": "campaign-chi-drift",
 "hypotheses": [
  {
   "id": "H-chi-bon",
   "prediction": {
    "ci_excludes": 0,
    "comparator": ">",
    "effect": null,
    "metric": "chi_bon_spearman",
    "threshold": 0.3
   },
   "scoreboard_row": "T9",
   "statement": "chi predicts realized best-of-n feature-drift ordering on a real bank"
  }
 ],
 "science": "S03-thermo",
 "registered_probability": 0.6
}

step 2 of 5

The specs are frozen

All 27 specs were hashed and tagged before the run. The +dirty suffix on the git sha is itself disclosed. Every threshold predates the confirmatory evidence by construction.

specs/frozen/manifest.json (fragment)

{
 "created_at": "2026-07-18T23:46:57.951556+00:00",
 "git_sha": "f93f4b5578c221d3d2665f9b305f3e86759be862",
 "tag": "campaign-freeze-v1",
 "spec_hashes": [
  "spec:c8ac36e72adcc96a3d65b802efbce423",
  "spec:39a99a5ab61d6f8e744a40b3804d21ab",
  "spec:59c457674b4b4f6fc20df702d3e34785",
  "… 24 more"
 ]
}

step 3 of 5

The fleet runs, on metered time

The confirmatory run cost $17.73 of metered H100 time across ten reward models. Every evidence record carries its own cost fields; this one is the single most expensive record in the store.

one evidence record from the merged store (value elided)

{
 "id": "ev:557a54240e62ac124ab41a1317cf01cb",
 "observable": "campaign.checkpoints.loadings",
 "subject": {
  "signals": [
   "skywork-v2-qwen3-0.6b"
  ],
  "dataset": null,
  "readout": null,
  "frame": null,
  "interventions": [],
  "extra": {
   "roster_key": "skywork-v2-qwen3-0.6b",
   "slice": "ultrafeedback-train-20k"
  }
 },
 "value": "<campaign.payloads.FeatureMatrix elided>",
 "trust": 0,
 "provenance": {
  "git_sha": "unknown",
  "cost": {
   "gpu_seconds": 1253.654161738,
   "tokens": 0,
   "wall_seconds": 1253.654161738
  },
  "study": null,
  "extra": {
   "gpu": "H100",
   "arc": "train:emb-lora"
  }
 },
 "created_at": "2026-07-19T15:21:07.925942+00:00"
}

step 4 of 5

Adjudication happens by machinery

The adjudicator compares each observed value against its frozen threshold and writes the verdict into the append-only store. CHI-DRIFT observed 0.9643 against its registered 0.3.

adjudication record ev:c978e3958769ecf93dcf079771a3aedb (fragment)

{
 "card": "CHI-DRIFT",
 "outcomes": {
  "H-chi-bon": "confirmed",
  "H-chi-beats-var": "confirmed"
 },
 "metrics": {
  "chi_bon_spearman": 0.9642857142857145,
  "chi_minus_variance_baseline_spearman": 0.5714285714285714
 },
 "killed": false,
 "evidence_id": "ev:c978e3958769ecf93dcf079771a3aedb"
}

step 5 of 5

Refutations are published with the same prominence

CAL-TRANSFER observed 0.419 against its frozen 0.15 equivalence bound. K-transfer fired. The refutation runs through exactly the same machinery as every confirmation, and it leads the results document.

adjudication record ev:b37f71c015ebfd697d7deb2a7efb5f7e (fragment)

{
 "card": "CAL-TRANSFER",
 "outcomes": {
  "H-transfer": "refuted"
 },
 "metrics": {
  "max_abs_auc_difference": 0.41898333333333326
 },
 "killed": true,
 "killed_by": [
  "K-transfer"
 ],
 "evidence_id": "ev:b37f71c015ebfd697d7deb2a7efb5f7e"
}

The hypothesis board

Every frozen hypothesis with its registered prediction, the observed value, and the outcome. Filters and sorting are linkable.

HypothesisRegistered prediction
ERR-RB2H-err-qweninternal features beat the exposed seen-pair margin at predicting held-out RB2 completion errors on Qwen3-8Berr_auroc_delta_qwen8b > 0.03p 0.6-9.74e-4 refutedK-err fired
ERR-RB2H-err-llamathe same holds on Llama-8B (cross-family)err_auroc_delta_llama8b > 0.03-0.003471 refutedK-err fired
FORECAST-CATH-forecast-maeinternals plus a 200-pair slice forecast per-category RB2 accuracy within 0.06 MAEforecast_mae < 0.06p 0.550.08696 refuted
FORECAST-CATH-forecast-beats-sizethe forecast beats predict-from-sizeforecast_mae_minus_size_baseline < 0-0.1024 confirmed
STYLE-RMBH-style-transferinternal style susceptibility predicts the fleet's RM-Bench Hard rankingspearman_biasbattery_vs_rmbench_hard > 0.6p 0.5 inconclusive
STYLE-RMBH-style-permthe fleet-level correlation clears its permutation nullbiasbattery_rmbench_perm_p < 0.05 inconclusive
STYLE-RMBH-style-baselinethe internal battery at least matches the behavioral degradation baselinespearman_minus_behavioral_baseline >= 0 inconclusive
PPE-BONH-tail-plateauthe tail index predicts where best-of-k stops helping on the five PPE verifiable setsspearman_tail_vs_bon_plateau > 0.3p 0.55 inconclusive
PPE-BONH-lowerquantilethe lower-quantile best-of-n aggregation beats the mean against human labels, reproducing the PPE findingbon_lowerquantile_beats_mean == 1 inconclusive
LADDERH-ladder-coverageat least 2 of 3 capture-only internal quantities land in their pre-frozen 8B intervalsladder_8b_metrics_in_interval >= 2p 0.653 confirmed
LADDERH-ladder-beats-flatthe extrapolation beats a no-trend (equal-to-4B) baselineladder_extrap_minus_flat_error < 00.6002 refuted
CHI-DRIFTH-chi-bonchi predicts realized best-of-n feature-drift ordering on a real bankchi_bon_spearman > 0.3p 0.60.9643 confirmed
CHI-DRIFTH-chi-beats-varchi beats a base-variance-only baselinechi_minus_variance_baseline_spearman > 00.5714 confirmed
HUMPH-hump-existsan interior gold-reward hump exists in the proxy/gold best-of-n armhump_interior_present == 1p 0.55 inconclusive
HUMPH-hump-locthe tail-predicted hump location matches the observed one within 0.5 natshump_kl_abs_error < 0.5 inconclusive
GAUGE-E19H-rawthe raw cross-fine-tune reward cosine reproduces the v1 fixture near 0.005raw_cos_v01_v02 abs< 0.02p 0.9 inconclusive
GAUGE-E19H-canonthe frame-fixed cosine is far larger than the raw cosine (co-rotation)canonical_minus_raw > 0.4 inconclusive
GAUGE-E19H-localizethe top residual-angle concept matches the largest behavioral delta (formality)residual_concept_matches_top_behavioral_delta == 1 inconclusive
GAUGE-XFAMH-xfamthe canonical angle between two RMs' reward directions predicts their held-out disagreement ratespearman_canonical_angle_vs_disagreement > 0.6p 0.5 inconclusive
GAUGE-XFAMH-xfam-permthe correlation clears a model-clustered permutation nullxfam_perm_p < 0.05 inconclusive
HACK-FOREH-flagweights-derived indices flag the realized-most-exploitable hack family above a behavioral baselineflag_hit_rate_minus_behavioral > 0p 0.55 inconclusive
SURGERYH-eraseLEACE erasure of the flagged direction removes the exploitexploit_drift_reduction > 0.8p 0.50.8856 confirmed
SURGERYH-retainbenchmark accuracy is retained through the erasurerb2_accuracy_delta_after_erasure > -0.01-0.3988 refuted
VERIF-PRMH-verif-locthe PRM localizes ProcessBench error stepsdense_localization_auc > 0.7p 0.60.2821 refutedK-verif fired
VERIF-PRMH-verif-stylethe correctness preference is anchored, not style-carriedstyle_share < 0.5-0.4268 confirmedK-verif fired
JUDGE-VBCH-vbcthe verdict decoded at the pre-critique position matches the final verdictverdict_prefix_match_rate >= 0.9p 0.61 confirmed
VALUES-CONTESTH-contestcontested-direction loading predicts per-item human rater disagreementcontested_raterspread_spearman > 0.2p 0.55 inconclusive
VALUES-CONTESTH-contest-vs-ensemblesingle-model contested loading comes within 0.05 of the 5x-cost ensemble baselineensemble_minus_contested_spearman < 0.05 inconclusive
FORENSIC-RECEIPTH-reliancethe two group contrasts separate graders as their planted values wouldreliance_recovery > 0.5p 0.60 refutedK-forensic fired
FORENSIC-RECEIPTH-credulitythe omission-credulity gap is real on constructed triplesomission_credulity_gap > 0.20.02326 refutedK-forensic fired
ATLAS-VCEH-vcevalue-convergence excess on real reward-model pairs beats the capability-matched random-utility nullreal_vce > 0p 0.5-0.06987 refutedK-vce fired
ATLAS-VCEH-vce-pthe convergent-pair alignment exceeds the random-utility null at p < 0.05reward_convergent_p_value < 0.050.001 confirmedK-vce fired
FACT-KUIH-kui-gapthe fleet-median KUI gap for the star property exceeds its permuted nullkui_gap_fleet_median > 0.2p 0.550.1515 refutedK-kui fired
FACT-KUIH-kui-permthe gap clears its permutation nullkui_gap_perm_p < 0.050.3886 refutedK-kui fired
FACT-KUIH-kui-controla priced control property shows no such gapkui_control_fleet_median < 0.10.3536 refutedK-kui fired
CAPACITY-WELCHH-dark-kdeffdark reward grows with the criteria-to-effective-dimension ratiodark_reward_kdeff_spearman > 0.6p 0.5-0.4062 refuted
CAPACITY-WELCHH-capacity-permthe fleet correlation clears its permutation nullcapacity_perm_p < 0.050.8934 refuted
CAPACITY-WELCHH-welchthe ArmoRM 19-head coherence matrix respects the Welch floorwelch_floor_min_slack >= 00.8164 confirmed
EVAL-AWAREH-aware-probea benchmark-versus-organic direction is decodablereal_probe_balanced_acc > 0.6p 0.60.7932 confirmed
EVAL-AWAREH-aware-steersteering along the direction inflates the scoredelta_r_per_steer > 0.05-2.10e-4 refuted
EVAL-AWAREH-aware-organic-fprthe probe does not spuriously flag organic prompts (equivalence bound, an affirmative null)organic_fpr < 0.10.317 refuted
TOPO-HODGEH-intransitivereal k-wise preference data carries a computable intransitive fractionintransitive_mass > 0.03p 0.70.214 confirmed
EMB-LORAH-bias-firstsurface biases reach half their final loading earlier than quality featuresreal_bias_before_quality > 1p 0.60.2692 refuted
EMB-LORAH-stabilizethe reward direction settles rather than wanderingw_r_stabilization_monotone == 11 confirmed
CAL-TRANSFERH-transferinstrument scorecard AUCs on a real 0.6B trunk match the CPU-organism AUCs within 0.15max_abs_auc_difference < 0.150.419 refutedK-transfer fired
ADJ-AVPH-avppatching's ranking matches the organism's marker-responsive components (the twin-activation-delta answer key) better than attribution'srecovery_gap > 0p 0.65-0.566 refutedK-avp fired
CONF-PARTIALH-partial-permthe surviving index clears its permutation nullbest_index_perm_p < 0.05p 0.50.0468 confirmed
CONF-PARTIALH-partialat least one load-bearing index retains a partial correlation above 0.3 after the size-and-accuracy controlmax_abs_partial_corr > 0.30.6485 confirmed
CONF-PARTIALH-partial-lineagethe surviving index also survives the within-lineage checklineage_check_pass == 10 refuted
META-LEDGERH-brierthe campaign's directional predictions, each assigned a pre-run probability, beat a coinbrier_directional < 0.250.26 refutedK-meta fired
META-LEDGERH-coveragethe nominal-0.8 prediction intervals cover the observed values near their nominal rateinterval_coverage_0p8 >= 0.60.75 confirmedK-meta fired
T3-DECOMPH-tacita residual no short rubric capturesreal_tacit_fraction > 00.9723 confirmed
T3-FIELDH-flatthe flat directions of the reward Hessian overlap the HACK-FORE-flagged direction more than a random direction doesflat_hack_overlap_minus_random > 0 inconclusive

The 27 cards

Confirmed 4

Refuted 8

Mixed 7

Inconclusive 8

The honest engineering note

Eight cards came back inconclusive for engineering reasons, not scientific ones: two hit a permission error in subject resolution, and six were missing an intermediate because an arc never ran or its shard never merged. The campaign runbook can re-run them against the same frozen specs. Eight other cards were frozen under explicit low-power acceptances, recorded in the freeze manifest before any evidence existed; each affected card page discloses its power and minimum detectable effect. The full accounting is on the method page.