Evidence records
Every number on this site came from somewhere.
Every number rendered anywhere on this site is bound to a record. A record carries the observable, the value with its interval, the trust level, the subject it was measured on, what it cost, and the commit that produced it. The binding is checked when the site is built: a number that resolves to neither a record nor a reasoned exception fails the build.
Find a record
Filter by what you remember about a number rather than by its identifier. The identifier is how the machine finds a record and it is the last thing anyone remembers.
- Susceptibility chi predicts best-of-n feature drift on a real bank skywork-v2-qwen3-8b, skywork-v2-llama31-8b c978e3958769ecf93dcf079771a3aedb
- Verdict before critique on a real generative judge skywork-critic e7fd3fee89a0c56cf4465212601d0d26
- Intransitive preference mass no scalar reward can represent nectar-tournaments, ultrafeedback-tournaments 844ad06a5aeb1d89c21c13c19a28cbd2
- A legibility frontier leaves a tacit residual skywork-v2-llama31-8b a7ed2d6b81c5f79458dad733defb3803
- Error anticipation on RewardBench 2 skywork-v2-qwen3-8b, skywork-v2-llama31-8b 6bf316c059230ff4ff99b118b1664631
- Does the process reward model verify or style-read qwen-prm 46e1500c4d39f9821c137ae0b00ac0da
- Skepticism and receipt reliance predict who rewards fabricated receipts fleet-subset 426a8cbed7f6d6dfb61d8fe3ed5264a4
- Value-convergence excess across the fleet beats the capability-matched null fleet 0c80d9086ebb0ad9704b88a8eefb01fa
- The fleet decodes a property it does not price fleet a34316feccee3beaed93c05ce5888567
- Instrument calibration transfers from a CPU organism to a real 0.6B trunk skywork-v2-qwen3-0.6b b37f71c015ebfd697d7deb2a7efb5f7e
- Attribution versus patching, resolved against planted ground truth skywork-v2-qwen3-0.6b, skywork-v02 a34457e83cd4eec10d080804a2533e82
- The v2 campaign scores its own calibration against a coin not stated 3990df7af1488c08ad62786d879b0b0f
- Benchmark forecasting from internals plus a 200-pair slice fleet 1acdc14e2fa16124786a9cd0d44a03b2
- Same-recipe scale ladder with frozen 8B extrapolation skywork-v2-qwen3-0.6b, skywork-v2-qwen3-1.7b, skywork-v2-qwen3-4b, skywork-v2-qwen3-8b d746d48eef7a5630cecbd3e97a9d9cb8
- Certified erasure removes the exploit without benchmark loss hackfore-flagged 41058731abf1d0b1590d4f2286e8771c
- Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor fleet, armorm f9e6a003a5a1e0aa0fe3ff1eca0d9c7b
- The RM knows it is being tested, and steering along that direction inflates reward evalaware-rms 54fb272fd011c8ca07a4959f553341ab
- Surface biases enter the reward direction before quality features skywork-v2-qwen3-0.6b 01c81c7185640f9b962b9115f360732e
- Do the indices carry information beyond size, accuracy, and lineage fleet 491f7f378d0b58dfdadebc9aaee7c983
- Bias battery predicts RM-Bench Hard ranking fleet 14bc4c666b5b115cf1fd2c86cf209f01
- Tail index forecasts the best-of-k plateau on PPE skywork-v2-qwen3-8b, skywork-v2-llama31-8b, skywork-v2-qwen3-0.6b 5cd71da5713f2b2fb77697318b842401
- The Goodhart hump located from a base-policy tail index policy-qwen25-1.5b, skywork-v2-qwen3-0.6b, skywork-v2-llama31-8b 02a1669641150750737afc8291255bf9
- The two-fine-tunes gauge result, settled skywork-v01, skywork-v02 0b9b34f7e2af83d3b0ca32a60791a859
- Canonical cross-family angles predict pairwise disagreement fleet 095096260d9209c4fb9935cbb215812b
- Weights-derived indices flag the exploitable hack family before any attack fleet 137077efe204a6daddd359f54c494453
- Contested-direction loading predicts human rater disagreement skywork-v2-llama31-8b, ensemble-5 e036326c47c7a0f52330559479e03790
- The reward Hessian's flat subspace overlaps the flagged hackable direction skywork-v2-llama31-8b 9d51086b170b359d58d8eb4bf4f882fb
Nothing matched those filters. Clear the subject, which is the one that most often returns nothing once a result is chosen beside it, or search for a word from the claim instead.
How much of this is covered
The build sweeps every numeral in the rendered text of every page and finds 5084 of them. 2419 resolve to a record. 2665 do not, and that is the larger number. Of those, 1427 sit on routes the sweep defers and does not check at all, 367 are on the generated API reference, where the figures are the library's own and this archive has no record to bind them to, 144 are a release version, an interpreter version, a figure or section label, a preprint address, a table's baseline row or a definition ported from the library's registry, each recognised by what surrounds it and checked against the file the build generated rather than trusted for its shape, and 3 forms of numeral are permanent exceptions with a stated reason: a date, a version string or an axis label is a legitimate exception and a measurement is not.
That is a count of numerals, not of distinct values: the same figure quoted on three pages counts three times, because what the build can check is occurrences. The exception list is published rather than hidden, and so is the shortfall, because a claim that every number resolves is worth nothing unless the count of the ones that do not is printed beside it rather than left to subtraction. The deferred routes are the honest part of this to look at: they are not numbers known to be unbound, they are numbers nobody has examined yet. Binding them is unfinished, and the site does not describe itself as though it were.
- ^(19|20|21|22|23|24|25|26)\d{2}$
- A four-digit year. A date is not a measurement, and every date on the site is either a commit date rendered from generated metadata or the year in a citation.
- ^[0-9]$
- A bare single digit. These are almost entirely list markers, heading ordinals and 'one of three' prose. A single digit cannot be a measurement with an uncertainty, and requiring a provenance chip on the '3' in 'three gates' would make the interface unreadable, which is the failure mode D-20 warns about for provenance as decoration.
- ^1?[0-9]{1,2}(px|rem|em|%|ch|vw|vh)$
- A CSS length that leaked into text. Not a claim about the world.
- /docs/**
- The ported documentation corpus, 74 routes with unbound numerics. Every one of these is a real unbound number today.
- /atlas/**
- Ten model pages whose measurements render as plain text.
- /evidence/**
- C-24: the record page is a JSON dump, so every field of every record renders as an unbound numeral.
- /log/**
- The v2 campaign log.
- /
- Its numbers are bound when it is rebuilt.
- /gallery/
- Every number on the component gallery is a fixture, which is exactly why D-33 has this route deleted before launch and why gate 5 would otherwise fail on it. The gallery-absent build assertion removes this entry by removing the route.