Is the objective drifting while the run is still going?
My policy is training right now. Is the objective it is actually optimizing drifting away from the one I wrote, and will I know before it is obvious?
What was measured, and with what access
A public labelled reinforcement-learning run with reward hacking: reward-lens-assay/PREDICTIONS.md:154 labelled rollouts across reward-lens-assay/PREDICTIONS.md:154 training steps, or reward-lens-assay/PREDICTIONS.md:154 by the run's other record, which is the count its own rollout table carries. The hacking transition is fitted at step reward-lens-assay/PREDICTIONS.md:154 with a 10-to-90 width of reward-lens-assay/PREDICTIONS.md:154 steps.
| labelled rollouts | reward-lens-assay/PREDICTIONS.md:154 |
|---|---|
| training steps, the lower of two records | reward-lens-assay/PREDICTIONS.md:154 |
| fitted transition midpoint | reward-lens-assay/PREDICTIONS.md:154 |
| 10-90 transition width | reward-lens-assay/PREDICTIONS.md:154 |
| fit quality, R squared | reward-lens-assay/PREDICTIONS.md:154 |
Move through the run and watch the signals
Read each row across before choosing one. The signal with the longest apparent lead is also the one that fires most often on order-destroyed surrogates of the same series, which is why it was not the one kept. Rows stay in the study's own evaluation order and are never sorted by lead.
| step | hack rate |
|---|---|
| 1 | 0.016 |
| 2 | 0.031 |
| 3 | 0.000 |
| 4 | 0.000 |
| 5 | 0.016 |
| 6 | 0.000 |
| 7 | 0.016 |
| 8 | 0.000 |
| 9 | 0.016 |
| 10 | 0.000 |
| 11 | 0.016 |
| 12 | 0.016 |
| 13 | 0.000 |
| 14 | 0.000 |
| 15 | 0.000 |
| 16 | 0.000 |
| 17 | 0.000 |
| 18 | 0.016 |
| 19 | 0.000 |
| 20 | 0.000 |
| 21 | 0.016 |
| 22 | 0.031 |
| 23 | 0.000 |
| 24 | 0.000 |
| 25 | 0.000 |
| 26 | 0.000 |
| 27 | 0.016 |
| 28 | 0.000 |
| 29 | 0.000 |
| 30 | 0.000 |
| 31 | 0.000 |
| 32 | 0.000 |
| 33 | 0.016 |
| 34 | 0.016 |
| 35 | 0.016 |
| 36 | 0.000 |
| 37 | 0.000 |
| 38 | 0.000 |
| 39 | 0.000 |
| 40 | 0.000 |
| 41 | 0.000 |
| 42 | 0.000 |
| 43 | 0.016 |
| 44 | 0.000 |
| 45 | 0.000 |
| 46 | 0.000 |
| 47 | 0.016 |
| 48 | 0.000 |
| 49 | 0.000 |
| 50 | 0.000 |
| 51 | 0.000 |
| 52 | 0.000 |
| 53 | 0.000 |
| 54 | 0.016 |
| 55 | 0.016 |
| 56 | 0.016 |
| 57 | 0.000 |
| 58 | 0.000 |
| 59 | 0.000 |
| 60 | 0.000 |
| 61 | 0.016 |
| 62 | 0.000 |
| 63 | 0.000 |
| 64 | 0.000 |
| 65 | 0.016 |
| 66 | 0.000 |
| 67 | 0.000 |
| 68 | 0.000 |
| 69 | 0.000 |
| 70 | 0.031 |
| 71 | 0.031 |
| 72 | 0.031 |
| 73 | 0.000 |
| 74 | 0.000 |
| 75 | 0.016 |
| 76 | 0.000 |
| 77 | 0.000 |
| 78 | 0.000 |
| 79 | 0.031 |
| 80 | 0.000 |
| 81 | 0.016 |
| 82 | 0.000 |
| 83 | 0.031 |
| 84 | 0.000 |
| 85 | 0.016 |
| 86 | 0.016 |
| 87 | 0.031 |
| 88 | 0.047 |
| 89 | 0.047 |
| 90 | 0.063 |
| 91 | 0.000 |
| 92 | 0.016 |
| 93 | 0.031 |
| 94 | 0.063 |
| 95 | 0.063 |
| 96 | 0.078 |
| 97 | 0.094 |
| 98 | 0.125 |
| 99 | 0.078 |
| 100 | 0.156 |
| 101 | 0.109 |
| 102 | 0.250 |
| 103 | 0.375 |
| 104 | 0.547 |
| 105 | 0.500 |
| 106 | 0.547 |
| 107 | 0.516 |
| 108 | 0.656 |
| 109 | 0.547 |
| 110 | 0.578 |
| 111 | 0.734 |
| 112 | 0.625 |
| 113 | 0.781 |
| 114 | 0.766 |
| 115 | 0.828 |
| 116 | 0.875 |
| 117 | 0.859 |
| 118 | 0.906 |
| 119 | 0.922 |
| 120 | 0.859 |
| 121 | 0.891 |
| 122 | 0.766 |
| 123 | 0.859 |
| 124 | 0.906 |
| 125 | 0.781 |
| 126 | 0.859 |
| 127 | 0.969 |
| 128 | 0.891 |
| 129 | 0.875 |
| 130 | 0.969 |
| 131 | 0.891 |
| 132 | 0.969 |
| 133 | 0.938 |
| 134 | 0.938 |
| 135 | 0.922 |
| 136 | 0.922 |
| 137 | 0.938 |
| 138 | 0.969 |
| 139 | 0.969 |
| 140 | 1.000 |
| 141 | 0.984 |
| 142 | 0.984 |
| 143 | 0.984 |
| 144 | 0.984 |
| 145 | 1.000 |
| 146 | 1.000 |
| 147 | 0.969 |
| 148 | 0.984 |
| 149 | 0.969 |
| 150 | 1.000 |
| 151 | 0.984 |
| 152 | 1.000 |
| 153 | 1.000 |
| 154 | 0.984 |
| 155 | 1.000 |
| 156 | 1.000 |
| 157 | 0.984 |
| 158 | 1.000 |
| 159 | 0.984 |
| 160 | 0.984 |
| 161 | 1.000 |
| 162 | 0.969 |
| 163 | 1.000 |
| 164 | 1.000 |
| 165 | 0.969 |
| 166 | 0.969 |
| 167 | 1.000 |
| 168 | 0.953 |
| 169 | 1.000 |
| 170 | 0.984 |
| 171 | 1.000 |
| 172 | 1.000 |
| 173 | 1.000 |
| 174 | 0.984 |
| 175 | 0.969 |
| 176 | 1.000 |
| 177 | 1.000 |
| 178 | 0.984 |
| 179 | 1.000 |
| 180 | 1.000 |
| 181 | 1.000 |
| 182 | 1.000 |
| 183 | 1.000 |
| 184 | 1.000 |
| 185 | 0.969 |
| 186 | 1.000 |
| 187 | 0.984 |
| 188 | 1.000 |
| 189 | 0.984 |
| 190 | 0.984 |
| 191 | 0.984 |
| 192 | 0.984 |
| 193 | 0.984 |
| 194 | 0.969 |
| 195 | 1.000 |
| 196 | 0.969 |
| 197 | 1.000 |
| 198 | 1.000 |
| 199 | 0.984 |
| 200 | 0.969 |
| 201 | 0.984 |
| 202 | 0.984 |
| 203 | 0.984 |
| 204 | 1.000 |
| 205 | 0.984 |
| 206 | 0.984 |
| 207 | 1.000 |
| 208 | 1.000 |
| 209 | 0.969 |
| 210 | 1.000 |
| 211 | 0.984 |
| 212 | 0.969 |
| 213 | 0.969 |
| 214 | 1.000 |
| 215 | 0.969 |
| 216 | 1.000 |
| 217 | 1.000 |
| 218 | 1.000 |
| 219 | 0.984 |
| 220 | 0.969 |
| 221 | 1.000 |
| 222 | 0.984 |
| 223 | 1.000 |
| 224 | 0.984 |
| 225 | 1.000 |
| 226 | 1.000 |
| 227 | 0.984 |
| 228 | 0.984 |
| 229 | 0.984 |
| 230 | 0.984 |
| 231 | 1.000 |
| 232 | 0.969 |
| 233 | 1.000 |
| 234 | 1.000 |
| 235 | 0.984 |
| 236 | 0.969 |
| 237 | 1.000 |
| 238 | 0.984 |
| 239 | 1.000 |
| 240 | 1.000 |
| 241 | 1.000 |
| 242 | 0.969 |
| 243 | 1.000 |
| 244 | 1.000 |
| 245 | 0.984 |
| 246 | 1.000 |
| 247 | 0.984 |
| 248 | 1.000 |
| 249 | 1.000 |
| 250 | 0.984 |
| 251 | 0.984 |
| 252 | 0.984 |
| 253 | 1.000 |
| 254 | 1.000 |
| 255 | 1.000 |
| 256 | 1.000 |
| 257 | 1.000 |
| 258 | 1.000 |
| 259 | 1.000 |
| 260 | 1.000 |
| 261 | 1.000 |
| 262 | 1.000 |
| 263 | 1.000 |
| 264 | 1.000 |
| 265 | 1.000 |
| 266 | 1.000 |
| 267 | 1.000 |
| 268 | 1.000 |
| 269 | 1.000 |
| 270 | 1.000 |
| 271 | 1.000 |
| 272 | 1.000 |
| 273 | 1.000 |
| 274 | 0.984 |
| 275 | 1.000 |
| 276 | 1.000 |
| 277 | 1.000 |
| 278 | 1.000 |
| 279 | 1.000 |
| 280 | 1.000 |
| 281 | 1.000 |
| 282 | 1.000 |
| 283 | 1.000 |
| 284 | 1.000 |
| 285 | 0.984 |
| 286 | 0.984 |
| 287 | 1.000 |
| 288 | 0.984 |
| 289 | 0.984 |
| 290 | 1.000 |
| 291 | 1.000 |
| 292 | 0.984 |
| 293 | 0.984 |
| 294 | 1.000 |
| 295 | 1.000 |
| 296 | 1.000 |
| 297 | 1.000 |
| 298 | 1.000 |
| 299 | 1.000 |
| 300 | 1.000 |
| 301 | 1.000 |
| 302 | 1.000 |
| 303 | 1.000 |
| 304 | 1.000 |
| 305 | 1.000 |
| 306 | 0.984 |
| 307 | 0.984 |
| 308 | 1.000 |
| 309 | 0.984 |
| 310 | 1.000 |
| 311 | 1.000 |
| 312 | 0.984 |
| 313 | 1.000 |
| 314 | 1.000 |
| 315 | 1.000 |
| 316 | 1.000 |
| 317 | 1.000 |
| 318 | 1.000 |
| 319 | 1.000 |
| 320 | 0.984 |
| 321 | 1.000 |
| 322 | 0.984 |
| 323 | 1.000 |
| 324 | 1.000 |
| 325 | 1.000 |
| 326 | 1.000 |
| 327 | 1.000 |
| 328 | 1.000 |
| 329 | 0.984 |
| 330 | 1.000 |
| 331 | 1.000 |
| 332 | 1.000 |
| 333 | 1.000 |
| 334 | 1.000 |
| 335 | 1.000 |
| 336 | 1.000 |
| 337 | 1.000 |
| 338 | 1.000 |
| 339 | 1.000 |
| 340 | 0.984 |
| 341 | 0.984 |
| 342 | 1.000 |
| 343 | 1.000 |
| 344 | 1.000 |
| 345 | 1.000 |
| 346 | 1.000 |
| 347 | 1.000 |
| 348 | 1.000 |
| 349 | 1.000 |
| 350 | 0.984 |
| 351 | 1.000 |
| 352 | 1.000 |
| 353 | 1.000 |
| 354 | 1.000 |
| 355 | 1.000 |
| 356 | 1.000 |
| 357 | 1.000 |
| 358 | 1.000 |
| 359 | 1.000 |
| 360 | 1.000 |
| 361 | 1.000 |
| 362 | 1.000 |
| 363 | 1.000 |
| 364 | 1.000 |
| 365 | 0.984 |
| 366 | 0.969 |
| 367 | 0.844 |
| 368 | 1.000 |
| 369 | 1.000 |
| 370 | 1.000 |
| 371 | 0.984 |
| 372 | 0.984 |
| 373 | 0.984 |
| 374 | 1.000 |
| 375 | 1.000 |
| 376 | 0.984 |
| 377 | 0.984 |
| 378 | 1.000 |
| 379 | 1.000 |
| 380 | 1.000 |
| 381 | 1.000 |
| 382 | 1.000 |
| 383 | 1.000 |
| 384 | 0.984 |
| 385 | 1.000 |
| 386 | 1.000 |
| 387 | 1.000 |
| 388 | 1.000 |
| 389 | 1.000 |
| 390 | 1.000 |
| 391 | 1.000 |
| 392 | 1.000 |
| 393 | 0.984 |
| 394 | 0.984 |
| 395 | 1.000 |
| 396 | 1.000 |
| 397 | 1.000 |
| 398 | 0.984 |
| 399 | 1.000 |
| 400 | 1.000 |
| 401 | 0.984 |
- step
- reward-lens-assay/experiments/x5_threshold/aisi_lengths.npz
- hack rate
- reward-lens-assay/experiments/x5_threshold/aisi_lengths.npz
- alarm
- reward-lens-assay/PREDICTIONS.md:154
- midpoint
- reward-lens-assay/PREDICTIONS.md:154
- comparator
- reward-lens-assay/PREDICTIONS.md:154
Two records of this run report its length differently, as reward-lens-assay/PREDICTIONS.md:154 steps and as reward-lens-assay/PREDICTIONS.md:154. This site uses reward-lens-assay/PREDICTIONS.md:154 and says so here rather than choosing one quietly. The public rollout table labels its steps from one, which accounts for the difference arithmetically without establishing which record is right.
One signal led. The comparator named in advance lagged.
The shipped signal leads. The comparator named in advance lags.
The variance derivative alarms at step reward-lens-assay/PREDICTIONS.md:154, a lead of reward-lens-assay/PREDICTIONS.md:154 transition widths. The gradient-norm comparator, named in advance, peaks at step reward-lens-assay/PREDICTIONS.md:154, which is reward-lens-assay/PREDICTIONS.md:154 steps after the midpoint, a lead of reward-lens-assay/PREDICTIONS.md:154 widths.
Two signals that lead by more, and why they were discarded.
The variance level alarms at step reward-lens-assay/PREDICTIONS.md:154 for a lead of reward-lens-assay/PREDICTIONS.md:154 widths and fires on reward-lens-assay/PREDICTIONS.md:154 of order-destroyed surrogates from the same series. The gradient norm under a cumulative-sum detector fires on reward-lens-assay/PREDICTIONS.md:154. The shipped derivative fires on reward-lens-assay/PREDICTIONS.md:154.
| detector | lead, transition widths | fires on surrogates |
|---|---|---|
| Within-group reward variance, level | reward-lens-assay/PREDICTIONS.md:154 | reward-lens-assay/PREDICTIONS.md:154 |
| Gradient norm under a cumulative-sum alarm | not measured | reward-lens-assay/PREDICTIONS.md:154 |
| Within-group reward variance, derivative shipped | reward-lens-assay/PREDICTIONS.md:154 | reward-lens-assay/PREDICTIONS.md:154 |
| Gradient-norm peak, the comparator named in advance | reward-lens-assay/PREDICTIONS.md:154 | not measured |
A surrogate keeps every value of the real series and destroys the order. A detector that fires as often on the surrogate as on the real run is reading the distribution, not the drift.
What this does not establish
This is one scalar series on one public run. No activations were read and no mechanism is named. The false-positive rate across runs is not settled, and the gradient-norm comparator could not be computed on the companion analysis at all, because the published rollout table carries no optimiser telemetry.
What this does not establish
| fires on surrogates, Within-group reward variance, derivative | reward-lens-assay/PREDICTIONS.md:154 |
|---|---|
| fires on surrogates, from the trainer log | reward-lens-assay/PREDICTIONS.md:154 |
| fires on surrogates, Within-group reward variance, level | reward-lens-assay/PREDICTIONS.md:154 |
| fires on surrogates, Gradient norm under a cumulative-sum alarm | reward-lens-assay/PREDICTIONS.md:154 |
| fires on surrogates, Gradient-norm peak, the comparator named in advance | not measured |
A defect this build found in its own frozen analysis.
The frozen path for a companion prediction standardised the series against a mean the post-transition regime had raised, so it used the future to define what normal looked like and reported a large lead that is an artifact. A detector that uses the future to define normal cannot measure a lead time, and it will report one. Both numbers are published, they disagree in sign, and the flattering one was declined.
The frozen path standardises the series against a mean the post-transition regime raised, so the accumulator crosses a threshold defined partly by the future. A detector that uses the future to define normal cannot measure a lead time, and it will report one. The detector-free comparison was added after the freeze, and it is labelled as such wherever it appears.
| the frozen path | reward-lens-assay/PREDICTIONS.md:136 |
|---|---|
| the corrected path | reward-lens-assay/PREDICTIONS.md:136 |
| flattering value declined | reward-lens-assay/PREDICTIONS.md:136 |
The companion prediction on the same run, which went the other way.
A second quantity, named in the same freeze, was predicted to move before the labelled hack rate. It lags, by reward-lens-assay/PREDICTIONS.md:155 transition widths. The sign is robust and the magnitude is not. What failed is the lead-time claim and not the instrument.
| Lambda moves before the labelled hack rate | reward-lens-assay/PREDICTIONS.md:155 |
|---|---|
| sign robust | reward-lens-assay/PREDICTIONS.md:155 |
| magnitude robust | reward-lens-assay/PREDICTIONS.md:155 |
- the frozen path
- reward-lens-assay/PREDICTIONS.md:136
- the corrected path
- reward-lens-assay/PREDICTIONS.md:136
The frozen path standardises the series against a mean the post-transition regime raised, so the accumulator crosses a threshold defined partly by the future. A detector that uses the future to define normal cannot measure a lead time, and it will report one. The detector-free comparison was added after the freeze, and it is labelled as such wherever it appears.
Run it yourself
unzip run.zip -d run && python run/reproduce.py Reproduce this
Everything on this page regenerates from the bundle below.
- contents
- PREDICTIONS-R3-R4.txt, aisi_lengths.npz, reproduce.py, run-series.json
- size
- public/bundles/run.zip
- sha256
-
ba9b5eba3ac2161016b5bdcec24d2a39cd6a7dff496792805f3438f7dd3852ad - commit
- public/bundles/run.zip
- needs
- Python 3.10 or newer. No GPU, no model and no network. The bundle carries the labelled rollout table, so the fit regenerates end to end from what is in it. The alarm does not: the rollout table records lengths and outcomes and carries no reward and no group index, so a within-group reward variance cannot be formed from it. The bundle re-fits the transition and prints its own numbers beside the ledger's.
Expected output
rollouts in the table 25664
rollouts in run-series.json 25664
distinct step indices 401 (1 to 401)
rollouts per step 64
step count the site states 400
series formatOk 401 points matches run-series.json: true
series hackRate 401 points matches run-series.json: true
series realized 401 points matches run-series.json: true
series written 401 points matches run-series.json: true
this fit frozen ledger
transition midpoint 106.9368 106.0000
10-to-90 width 23.9823 23.9000
R2 0.99555 0.99600
floor and ceiling 0.0022 to 0.9892
detector alarms, from the frozen ledger row in this bundle, not refitted here
variance_level step 70 lead +1.504 widths
gradient_norm_cusum no alarm
variance_derivative step 90 lead +0.668 widths <- shipped
gradient_norm_peak step 133 lead -1.129 widths No independent reproduction recorded.