Reward thermodynamics

Tail index forecasts the best-of-k plateau on PPE

inconclusive registered

The reading

  1. the tail index predicts where best-of-k stops helping on the five PPE verifiable sets

    spearman_tail_vs_bon_plateau > 0.3 0.55 not measured inconclusive

  2. the lower-quantile best-of-n aggregation beats the mean against human labels, reproducing the PPE finding

    bon_lowerquantile_beats_mean == 1 not measured inconclusive

PPE-BON was not adjudicated this run: missing intermediate: no intermediate 'campaign.scores' for roster_key='skywork-v2-qwen3-8b' slice='ppe-best-of-k::gpqa'; the arc that produces it has not run or its shard was not merged.
PPE-BON was not adjudicated this run: missing intermediate: no intermediate 'campaign.scores' for roster_key='skywork-v2-qwen3-8b' slice='ppe-best-of-k::gpqa'; the arc that produces it has not run or its shard was not merged.

Provenance

card
PPE-BON
frozen_at
adjudicated_at
spec_hash
spec:7077cf849ed8492fc327d970aa2ede55
study_id
study:campaign-ppe-bon@v1#7077cf84
power_acceptance
Fisher-z detection of the frozen 0.3 against the null-critical 0.46 over n = 15 (model, set) points; the permutation null passes the frozen 0.3 0.17 of the time, the registered CI exclusion carries size control, and effects at or above 0.64 confirm at 80 percent