The failed measurement
The first Stage-0 run tested whether GPT-2 small behaves like an iterative settling process. It measured relative residual updates across 12 layers and defined settling depth as the first layer whose top prediction matched the final layer's prediction and stayed matched.
Every easy and hard prompt returned a settling depth of 12. The tempting interpretation was that nothing settled early. The definition made that conclusion unavailable. A large rewrite in layer 12 changed the top prediction, so the only layer guaranteed to equal the final layer and remain equal was layer 12 itself. The metric measured its own stopping rule.
The vector of mean relative updates still contained useful information: [0.49, 1.06, 4.45, 0.42, 0.36, 0.27, 0.21, 0.20, 0.20, 0.23, 0.53, 7.04]. GPT-2 small showed a quiet middle and large rewrites near the boundaries. The run did not establish monotone decay or different settling depth for easy and hard prompts.
The repaired metric failed differently
Stage-0b replaced exact top-one agreement with KL divergence from each layer's logit-lens distribution to the final distribution. The threshold was 0.1 bits. Across nine prompts, the result was again 12 every time. Intermediate layers remained two to eight bits from the final distribution, then the final point dropped to zero by definition.
This second result is stronger than the first because the curve itself shows a late readout transition. The scalar summary is still saturated. A useful repair would normalize divergence to the initial layer, exclude the trivial final point, report the full curve, and compare a raw logit lens with a tuned lens.
A second bug in the same run
The lab's novelty index computed Shannon entropy over raw feature magnitudes. Adding a parameter-count field of 124,439,808 swamped the other features and drove the run's entropy to 0.01 bits. That was a scale bug, not a discovery about novelty. Applying log1p before normalization changed the indexed values to 8.25 bits for Stage-0 and 7.86 bits for Stage-0b.
This is why measurement software belongs inside the research result. The model run surfaced facts about GPT-2, but it also falsified two parts of the lab's own instrumentation. Publishing only the model curve would hide the more transferable lesson.
What remains defensible
GPT-2 small used 1,204 MiB peak RSS and 237 MiB of MPS allocation in Stage-0b. The prompt "The capital of France is" produced " Paris" only five times in 100 sampled first tokens; the fixed next-token distribution preferred grammatical continuations such as " the" and " a". Hidden states were deterministic for the fixed input, so the observed 100-run variation came from sampling one distribution.
The finding is limited to GPT-2 small, its tokenizer, this prompt position, and the raw logit-lens setup. It is not evidence that factual knowledge is absent from later positions or larger models. It is evidence that the immediate distribution at this scale is grammatical before it is encyclopedic.
Evidence ledger
Executed: Stage-0 and Stage-0b GPT-2 small runs, including memory measurements and layerwise curves.
Read:
settling-field-lab/interp/RESULTS.okf.mdandsettling-field-lab/interp/RESULTS-0b.okf.md.Not established: a general settling-depth law, an easy-versus-hard separation, or behavior in larger models.
Next falsifier: exclude the final point, normalize the curve, and compare raw and tuned lenses.



