Every number below is a view over the campaign's evidence store. The per-model captures are exploratory by construction (the calibration gates were exercised by the campaign cards, not by these fleet sweeps), and the trust badges say so rather than hiding it.

Model Weights crystallization snr Records GPU-seconds
skywork-v2-qwen3-0.6b Skywork/Skywork-Reward-V2-Qwen3-0.6B 0.895 E Exploratory 1.14 E Exploratory 290 2,746
skywork-v2-qwen3-1.7b Skywork/Skywork-Reward-V2-Qwen3-1.7B 0.93 E Exploratory 1.42 E Exploratory 55 173
skywork-v2-qwen3-4b Skywork/Skywork-Reward-V2-Qwen3-4B 0.923 E Exploratory 1.55 E Exploratory 55 370
skywork-v2-qwen3-8b Skywork/Skywork-Reward-V2-Qwen3-8B 0.916 E Exploratory 1.66 E Exploratory 186 2,860
skywork-v2-llama31-8b Skywork/Skywork-Reward-V2-Llama-3.1-8B 0.914 E Exploratory 1.73 E Exploratory 299 4,089
skywork-v01 Skywork/Skywork-Reward-Llama-3.1-8B 0.905 E Exploratory 1.85 E Exploratory 56 454
skywork-v02 Skywork/Skywork-Reward-Llama-3.1-8B-v0.2 0.904 E Exploratory 1.66 E Exploratory 77 868
armorm RLHFlow/ArmoRM-Llama3-8B-v0.1 0.554 E Exploratory 1.06 E Exploratory 68 567
tulu-rm allenai/Llama-3.1-Tulu-3-8B-RM 0.841 E Exploratory 0.906 E Exploratory 90 892
grm-gemma2-2b Ray2333/GRM-gemma2-2B-rewardmodel-ft 0.77 E Exploratory 1.5 E Exploratory 55 205

crystallization and snr come from each model's campaign.index.table on the rb2-full slice; models without a merged index table show a dash. The column headers link the docs that define each instrument, and the index library defines the rest of the per-model metrics. The per-model pages carry the full observable tables with evidence ids.