The archive

Value-convergence excess across the fleet beats the capability-matched null

refuted registered

The reading

  1. value-convergence excess on real reward-model pairs beats the capability-matched random-utility null

    real_vce > 0 0.5 Value-convergence excess across the fleet beats the capability-matched null refuted registered confidence interval -0.12281662932807336 to -0.009192229234832329 > 0 The frozen study "Value-convergence excess across the fleet beats the capability-matched null", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-helpfulness-512. It does not establish that the same holds on another subject, another dataset or another run. the vce > 0 margin sits at the null center by construction: the point call passes under the null 0.48 of the time and its power at the threshold-alternative equals that rate; the alpha-exact p-value companion and the registered CI exclusion carry the card's evidence. Open the record › refuted

  2. the convergent-pair alignment exceeds the random-utility null at p < 0.05

    reward_convergent_p_value < 0.05 Value-convergence excess across the fleet beats the capability-matched null refuted registered < 0.05 The frozen study "Value-convergence excess across the fleet beats the capability-matched null", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-helpfulness-512. It does not establish that the same holds on another subject, another dataset or another run. the vce > 0 margin sits at the null center by construction: the point call passes under the null 0.48 of the time and its power at the threshold-alternative equals that rate; the alpha-exact p-value companion and the registered CI exclusion carry the card's evidence. Open the record › confirmed

convergence is fully explained by shared capability structure

real_vce <= 0 Value-convergence excess across the fleet beats the capability-matched null refuted registered confidence interval -0.12281662932807336 to -0.009192229234832329 > 0 The frozen study "Value-convergence excess across the fleet beats the capability-matched null", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-helpfulness-512. It does not establish that the same holds on another subject, another dataset or another run. the vce > 0 margin sits at the null center by construction: the point call passes under the null 0.48 of the time and its power at the threshold-alternative equals that rate; the alpha-exact p-value companion and the registered CI exclusion carry the card's evidence. Open the record ›

ATLAS-VCE fired its kill criterion (K-vce); the registered negative result.
ATLAS-VCE fired its kill criterion (K-vce); the registered negative result.

Measurements

real_vce_ci_low
Value-convergence excess across the fleet beats the capability-matched null refuted registered confidence interval -0.12281662932807336 to -0.009192229234832329 > 0 The frozen study "Value-convergence excess across the fleet beats the capability-matched null", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-helpfulness-512. It does not establish that the same holds on another subject, another dataset or another run. the vce > 0 margin sits at the null center by construction: the point call passes under the null 0.48 of the time and its power at the threshold-alternative equals that rate; the alpha-exact p-value companion and the registered CI exclusion carry the card's evidence. Open the record ›
real_vce_ci_high
Value-convergence excess across the fleet beats the capability-matched null refuted registered confidence interval -0.12281662932807336 to -0.009192229234832329 > 0 The frozen study "Value-convergence excess across the fleet beats the capability-matched null", adjudicated on 19 July 2026 at the registered level of the project's evidence ladder. One frozen study, adjudicated once, on fleet, over rb2-full and rb2-helpfulness-512. It does not establish that the same holds on another subject, another dataset or another run. the vce > 0 margin sits at the null center by construction: the point call passes under the null 0.48 of the time and its power at the threshold-alternative equals that rate; the alpha-exact p-value companion and the registered CI exclusion carry the card's evidence. Open the record ›

Provenance

card
ATLAS-VCE
frozen_at
adjudicated_at
spec_hash
spec:2008bee7d02b9461abfd6f12ff7b0ffe
study_id
study:campaign-atlas-vce@v1#2008bee7
power_acceptance
the vce > 0 margin sits at the null center by construction: the point call passes under the null 0.48 of the time and its power at the threshold-alternative equals that rate; the alpha-exact p-value companion and the registered CI exclusion carry the card's evidence