Articles

Write-ups on our research, our open-source projects, and what we learn building them.

Follow along with the RSS feed.

research, cox, evaluation, benchmarks

Seven decision models on one test suite

Cox, with two readout heads, and six openly available decision models, evaluated on the same 21 tasks with the same corrected labels and the same scoring, and compared within three size classes. On the tasks that none of the models trained on, Cox 4B with its pointer head scores highest among the models of 4B parameters or fewer, a small but statistically reliable lead over Kev 4B. Kev 9B, at more than twice the size, scores higher still.

research, cox, interpretability

Where answers form inside a small decision model

We read the answer of a 4B decision model out of every layer of its backbone, for every kind of question it answers. Answers form in the middle of the network and the top third adds little; a hybrid backbone forms them earlier than a full-attention one; averaged models form them in the same layers as their members; answers that end up wrong settle later, though the model's own confidence already captures most of that; and with its fine-tuning switched off, the base backbone already holds much of the answer, which fine-tuning sharpens in the middle layers.

research, cox, evaluation

Mapping the question space: the task-prime matrix

Before training a decision model, we asked what kinds of question can be asked about a piece of text at all. The answer is a small grid, and it changed how we build data and how we measure.