A decision model reads a text and answers typed questions about it: yes or no, a choice among options, or a rating on a scale. Cox, our 4-billion-parameter decision model, does this by running the text and each question through a language-model backbone and reading the answer from the backbone’s final layer. This article asks a simple question about that design: at which layer does the answer actually exist?
To find out, we read the answer from every second layer of the backbone, for every kind of question in our test suite, in six models, and in one of them again with its fine-tuning switched off. Five findings:
- Answers form in the middle of the network. Nothing is readable in the first quarter; answers form between layers 8 and 20 of 32; the last third adds little on average.
- A hybrid backbone forms answers earlier than a backbone with full attention in every layer: the model’s own readout starts working at about 30% of the depth, against about 45%.
- Averaging trained models does not move the answers. A weight-averaged model (“soup”) forms its answers in the same layers as the models it was made from.
- Wrong answers settle later than right ones, but this adds little to what the model’s own confidence already says about its likely errors.
- Fine-tuning sharpens what the base model already has. With the fine-tuning switched off, the backbone’s middle layers already hold about three quarters of the answer; fine-tuning adds most in layers 14 to 20 and stops a decline in the top layers.
1. How we read an answer out of a layer
Cox uses a pointer readout. After the text and the question, the model writes an <answer> marker; each option is
also written out in the context, and the model scores each option by comparing the vector at the <answer> marker
with the vector at that option. Yes/no questions use a separate, simpler readout against two learned vectors.
Every layer of the backbone produces these vectors, not only the last. We read them out in two ways:
- A fresh readout per layer. For each layer we train a new pointer readout on that layer’s vectors (3,924 questions from the training data, the backbone frozen; three independently trained readouts per layer, averaged) and fit its temperature on a held-out part. This measures what the layer knows.
- The model’s own readout. We apply the trained model’s final readout, unchanged, to each layer, after matching the layer’s per-dimension statistics to those of the final layer. This measures when the model’s actual answer becomes readable. (It covers choice and rating questions; yes/no questions use the other readout.)
Both are scored on 400 questions from each of the 21 tasks in our evaluation suite (10,721 questions in all), the same questions for every model. The score is skill: 0 means no better than always choosing the most common answer, 1 means always right.
2. Answers form in the middle of the network
The main model here is our current best: a weight average of three training runs on Qwen3.5-4B.

| Layer read | 6 | 8 | 10 | 12 | 14 | 16 | 18 | 20 | 24 | 28 | 32 (final) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Fresh readout | +0.04 | +0.06 | +0.24 | +0.22 | +0.38 | +0.49 | +0.59 | +0.54 | +0.62 | +0.61 | +0.59 |
| The model’s own readout | −0.11 | −0.01 | +0.15 | +0.28 | +0.37 | +0.47 | +0.54 | +0.56 | +0.57 | +0.57 | +0.59 |
- Nothing is readable in the first quarter. Through layer 8, both readouts are at or near chance.
- Answers form between layers 8 and 20. The kinds of question differ a little: ratings on a scale level off by about layer 16 (+0.42 of a final +0.46); choices and yes/no answers keep improving to about layer 18–24.
- The last third adds little. The model’s own readout already scores +0.56 at layer 20 and +0.57 at layer 24, against +0.59 at the top. A fresh readout trained on layer 24 scores as well as one trained on the final layer.
3. Hybrid and full-attention backbones
Qwen3.5-4B is a hybrid: three of every four layers use a form of linear attention, which carries information forward in a fixed-size summary, and every fourth layer uses full attention, which can look back at every earlier token directly. To see what is specific to that design, we drew the same map for a Cox model on Qwen3-4B, which uses full attention in every one of its 36 layers. Because the two backbones differ in depth, the chart compares them by the share of the network read.

- The hybrid’s answers become readable earlier. Its own readout starts to work at about 30% of the depth and is at +0.47 by 50%. The full-attention model’s own readout is at chance until about 45% of its depth (+0.04 at layer 16 of 36), then rises steeply to +0.54 by 61%.
- Both finish at about 60% of the depth and are flat from there.
- With a fresh readout per layer the gap is smaller: the hybrid leads by roughly 5–10% of the depth.
We had expected steps right after the full-attention layers of the hybrid. One kind of question shows such a step in one model, but it is not a general pattern across those layers.
4. Averaged models form answers in the same layers
Our best models are soups: the weights of several training runs, averaged. We drew the map for two training runs of the same recipe (seeds 0 and 1), a third run trained with its readout first (“head first”), the average of the two seeds, and the average of all three, on the same questions.

- The curves rise in the same layers and end together. With a fresh readout per layer, all five models climb between layers 8 and 20 and are within about 0.03 of each other from layer 24 upward. In the middle, single points differ by up to 0.1 without a consistent order between models, which is within the variation of a trained readout. Averaging changes the final score a little, not where the answer is formed.
- The head-first run is readable a few layers earlier with its own readout (+0.21 at layer 10, against +0.09 to +0.15 for the others), and ends in the same place. Training the readout before the backbone appears to let the backbone’s earlier layers align with it sooner.
5. Wrong answers settle later
For each question we recorded the settling layer: the first layer from which the readout’s answer never changes again, all the way to the top.

- Questions the model answers right settle at a median of layer 12.
- Questions it answers wrong settle at a median of layer 18, and a quarter of them only at layer 30 or later (against 9% of the questions it answers right).
That suggests a way to flag likely errors. The comparison that matters is the model’s own confidence, the probability it gives its answer, which already flags errors. Measured as AUROC (how often a wrong answer is ranked as more doubtful than a right one):
| Model | Confidence | Settling layer | Both combined |
|---|---|---|---|
| seed 0 | 0.794 | 0.718 | 0.800 |
| seed 1 | 0.787 | 0.705 | 0.788 |
| head first | 0.784 | 0.692 | 0.789 |
| average of seeds 0 and 1 | 0.765 | 0.705 | 0.775 |
| average of all three | 0.759 | 0.692 | 0.765 |
| Qwen3-4B, full attention | 0.773 | 0.721 | 0.782 |
The settling layer is a weaker signal than confidence on its own, and adding it improves the combination by at most 0.010. The effect is real but mostly already captured by the final probability.
A second observation from the same table: the averaged models rank their own errors less well (0.76–0.77 against 0.78–0.79 for single runs), although their overall calibration is the same. An averaged model’s confidence separates its right and wrong answers less sharply, which matters wherever a model is asked to abstain when unsure.
6. What fine-tuning changes
Cox is trained by fine-tuning the backbone with small low-rank adapters (LoRA) in every layer, alongside the readout. To see what that fine-tuning does inside the network, we drew the map again for one training run (seed 0) with its adapters switched off: the original Qwen3.5-4B backbone, read the same two ways.

| Layer read | 8 | 12 | 14 | 16 | 18 | 20 | 24 | 32 (final) |
|---|---|---|---|---|---|---|---|---|
| Fresh readout, adapter off | +0.06 | +0.12 | +0.20 | +0.21 | +0.45 | +0.42 | +0.43 | +0.39 |
| Fresh readout, adapter on | +0.08 | +0.17 | +0.39 | +0.49 | +0.55 | +0.60 | +0.61 | +0.59 |
- The base backbone already holds much of the answer. A readout trained on its layer 18 reaches +0.45, about three quarters of the fine-tuned model’s best (+0.61), without any change to the backbone. Most of what a decision needs is already in the base model’s representations.
- Fine-tuning adds most in the middle. The gap between the two curves opens between layers 12 and 20, where answers form.
- It also stops a decline at the top. Without the adapters, readability peaks at layer 18 and falls to +0.39 at the final layer; the base model’s last layers turn towards predicting the next word, not deciding. With the adapters, the final layers keep the answer.
- The trained readout depends on the adapted vectors. Our own readout, applied to the base backbone, reaches only +0.23 at the final layer, against +0.59 with the adapters on.
A second test switched the adapters off in one group of four layers at a time, with the rest left on. No single group mattered much: the largest loss, from the first four layers, was 0.03, and every other group was within 0.015 of the full model. The adapters’ effect is spread across many layers, each partly able to stand in for the others.
7. A cheaper model? (future work)
If the answer is essentially complete by layer 20 or 24, the remaining 8 to 12 layers (25 to 37% of the backbone) could be removed, with the readout trained to read from the earlier layer. In an earlier experiment on a smaller, non-hybrid model (Qwen3-1.7B), doing so kept accuracy on familiar kinds of question but lost some on unfamiliar ones. We have not yet trained a truncated 4B model; until we have, the numbers above show what the layers contain, not that a shorter model would match the full one.
8. Limitations
- One model family and one suite. The maps are for Cox models on Qwen3.5-4B, with one Qwen3-4B contrast, scored on our own 21-task suite. Other backbones, readouts and tasks may place answers elsewhere.
- Readouts are trained. A fresh readout per layer shows what can be read from a layer with some training; the model’s own readout depends on matching each layer’s statistics to the final layer’s.
- Every second layer. Changes that happen between two measured layers are placed at the next one.
- Settling uses separately trained readouts per layer, so part of an answer’s change from layer to layer is the readouts disagreeing rather than the model changing its mind.
- Probabilities from the fresh readouts need large temperatures (about 10 to 30) to be calibrated, and most questions have two options, so per-question probability curves sit close to 0.5 until late.
- Variation between training runs: the two seeds differ by about 0.02 in final skill; differences between curves smaller than that should not be read into.
- The adapter comparison is one training run, read with its adapters on and off.
Method details
- Models: Cox pointer models on Qwen3.5-4B-Base with LoRA (rank 16), the half-length recipe (2,000 optimizer steps of 64 examples, cosine decay), trained on our research data mix; seeds 0 and 1, one run trained with the readout first (200 steps), and the averages of two and three runs with fitted temperatures. Contrast: a Cox pointer model on Qwen3-4B-Base (full attention). Adapters switched off: every LoRA update set to zero in seed 0 (the original backbone with the trained readout); for the layer-group test, in four layers at a time, scored on 120 questions per task.
- Vectors are taken from the model’s own pointer path, so they are exactly what its readout sees.
- Fresh readouts: a 256-dimensional bilinear pointer with learned yes/no vectors, 5 epochs, three per layer, temperature on 654 further held-out training questions.
- Error signals: AUROC against the model’s own errors; the combination is a two-feature logistic regression fitted out of fold (five folds).
- The analysis code is not yet public.