Find a record

Filter by what you remember about a number rather than by its identifier. The identifier is how the machine finds a record and it is the last thing anyone remembers.

How it came out
What it was measured on
  1. Susceptibility chi predicts best-of-n feature drift on a real bank skywork-v2-qwen3-8b, skywork-v2-llama31-8b confirmed registered c978e3958769ecf93dcf079771a3aedb
  2. Verdict before critique on a real generative judge skywork-critic confirmed registered e7fd3fee89a0c56cf4465212601d0d26
  3. Intransitive preference mass no scalar reward can represent nectar-tournaments, ultrafeedback-tournaments confirmed Withdrawn registered 844ad06a5aeb1d89c21c13c19a28cbd2
  4. A legibility frontier leaves a tacit residual skywork-v2-llama31-8b confirmed registered a7ed2d6b81c5f79458dad733defb3803
  5. Error anticipation on RewardBench 2 skywork-v2-qwen3-8b, skywork-v2-llama31-8b refuted registered 6bf316c059230ff4ff99b118b1664631
  6. Does the process reward model verify or style-read qwen-prm refuted registered 46e1500c4d39f9821c137ae0b00ac0da
  7. Skepticism and receipt reliance predict who rewards fabricated receipts fleet-subset refuted registered 426a8cbed7f6d6dfb61d8fe3ed5264a4
  8. Value-convergence excess across the fleet beats the capability-matched null fleet refuted registered 0c80d9086ebb0ad9704b88a8eefb01fa
  9. The fleet decodes a property it does not price fleet refuted registered a34316feccee3beaed93c05ce5888567
  10. Instrument calibration transfers from a CPU organism to a real 0.6B trunk skywork-v2-qwen3-0.6b refuted registered b37f71c015ebfd697d7deb2a7efb5f7e
  11. Attribution versus patching, resolved against planted ground truth skywork-v2-qwen3-0.6b, skywork-v02 refuted registered a34457e83cd4eec10d080804a2533e82
  12. The v2 campaign scores its own calibration against a coin not stated refuted registered 3990df7af1488c08ad62786d879b0b0f
  13. Benchmark forecasting from internals plus a 200-pair slice fleet mixed registered 1acdc14e2fa16124786a9cd0d44a03b2
  14. Same-recipe scale ladder with frozen 8B extrapolation skywork-v2-qwen3-0.6b, skywork-v2-qwen3-1.7b, skywork-v2-qwen3-4b, skywork-v2-qwen3-8b mixed registered d746d48eef7a5630cecbd3e97a9d9cb8
  15. Certified erasure removes the exploit without benchmark loss hackfore-flagged mixed registered 41058731abf1d0b1590d4f2286e8771c
  16. Dark reward tracks the criteria-to-dimension ratio; ArmoRM respects the Welch floor fleet, armorm mixed registered f9e6a003a5a1e0aa0fe3ff1eca0d9c7b
  17. The RM knows it is being tested, and steering along that direction inflates reward evalaware-rms mixed registered 54fb272fd011c8ca07a4959f553341ab
  18. Surface biases enter the reward direction before quality features skywork-v2-qwen3-0.6b mixed registered 01c81c7185640f9b962b9115f360732e
  19. Do the indices carry information beyond size, accuracy, and lineage fleet mixed registered 491f7f378d0b58dfdadebc9aaee7c983
  20. Bias battery predicts RM-Bench Hard ranking fleet inconclusive registered 14bc4c666b5b115cf1fd2c86cf209f01
  21. Tail index forecasts the best-of-k plateau on PPE skywork-v2-qwen3-8b, skywork-v2-llama31-8b, skywork-v2-qwen3-0.6b inconclusive registered 5cd71da5713f2b2fb77697318b842401
  22. The Goodhart hump located from a base-policy tail index policy-qwen25-1.5b, skywork-v2-qwen3-0.6b, skywork-v2-llama31-8b inconclusive registered 02a1669641150750737afc8291255bf9
  23. The two-fine-tunes gauge result, settled skywork-v01, skywork-v02 inconclusive registered 0b9b34f7e2af83d3b0ca32a60791a859
  24. Canonical cross-family angles predict pairwise disagreement fleet inconclusive registered 095096260d9209c4fb9935cbb215812b
  25. Weights-derived indices flag the exploitable hack family before any attack fleet inconclusive registered 137077efe204a6daddd359f54c494453
  26. Contested-direction loading predicts human rater disagreement skywork-v2-llama31-8b, ensemble-5 inconclusive registered e036326c47c7a0f52330559479e03790
  27. The reward Hessian's flat subspace overlaps the flagged hackable direction skywork-v2-llama31-8b inconclusive registered 9d51086b170b359d58d8eb4bf4f882fb

Nothing matched those filters. Clear the subject, which is the one that most often returns nothing once a result is chosen beside it, or search for a word from the claim instead.

How much of this is covered

The build sweeps every numeral in the rendered text of every page and finds 5084 of them. 2419 resolve to a record. 2665 do not, and that is the larger number. Of those, 1427 sit on routes the sweep defers and does not check at all, 367 are on the generated API reference, where the figures are the library's own and this archive has no record to bind them to, 144 are a release version, an interpreter version, a figure or section label, a preprint address, a table's baseline row or a definition ported from the library's registry, each recognised by what surrounds it and checked against the file the build generated rather than trusted for its shape, and 3 forms of numeral are permanent exceptions with a stated reason: a date, a version string or an axis label is a legitimate exception and a measurement is not.

That is a count of numerals, not of distinct values: the same figure quoted on three pages counts three times, because what the build can check is occurrences. The exception list is published rather than hidden, and so is the shortfall, because a claim that every number resolves is worth nothing unless the count of the ones that do not is printed beside it rather than left to subtraction. The deferred routes are the honest part of this to look at: they are not numbers known to be unbound, they are numbers nobody has examined yet. Binding them is unfinished, and the site does not describe itself as though it were.

^(19|20|21|22|23|24|25|26)\d{2}$
A four-digit year. A date is not a measurement, and every date on the site is either a commit date rendered from generated metadata or the year in a citation.
^[0-9]$
A bare single digit. These are almost entirely list markers, heading ordinals and 'one of three' prose. A single digit cannot be a measurement with an uncertainty, and requiring a provenance chip on the '3' in 'three gates' would make the interface unreadable, which is the failure mode D-20 warns about for provenance as decoration.
^1?[0-9]{1,2}(px|rem|em|%|ch|vw|vh)$
A CSS length that leaked into text. Not a claim about the world.
/docs/**
The ported documentation corpus, 74 routes with unbound numerics. Every one of these is a real unbound number today.
/atlas/**
Ten model pages whose measurements render as plain text.
/evidence/**
C-24: the record page is a JSON dump, so every field of every record renders as an unbound numeral.
/log/**
The v2 campaign log.
/
Its numbers are bound when it is rebuilt.
/gallery/
Every number on the component gallery is a fixture, which is exactly why D-33 has this route deleted before launch and why gate 5 would otherwise fail on it. The gallery-absent build assertion removes this entry by removing the route.