Certified robustness, attack surface and eval-awareness
Bias battery predicts RM-Bench Hard ranking
inconclusive registered
The reading
-
internal style susceptibility predicts the fleet's RM-Bench Hard ranking
spearman_biasbattery_vs_rmbench_hard > 0.6 0.5 not measured inconclusive
-
the fleet-level correlation clears its permutation null
biasbattery_rmbench_perm_p < 0.05 not measured inconclusive
-
the internal battery at least matches the behavioral degradation baseline
spearman_minus_behavioral_baseline >= 0 not measured inconclusive
Provenance
- card
- STYLE-RMB
- frozen_at
- adjudicated_at
- spec_hash
- spec:59c457674b4b4f6fc20df702d3e34785
- study_id
- study:campaign-style-rmb@v1#59c45767
- Evidence records
- 14bc4c666b5b115cf1fd2c86cf209f01
- spec_file
- campaign-style-rmb.json