Research

Cox

A small model that answers typed questions about text with calibrated probabilities, in one fast pass.

Software that reads text keeps needing small decisions. Which team should handle this ticket? Is this alert urgent? Do these two payment records match? Does this expense break a policy? A large language model can answer, but it replies in prose you have to parse, it takes seconds, and its “I’m fairly sure” is a sentence rather than a number you can act on.

Cox is our research into a better fit for those decisions.

Why we built it

Cox started earlier this year as a classifier for Vitund Sandbox’s events. A sandbox that records everything an agent does produces a stream of events, and many of them call for a quick judgement: is this request sending out data the agent read earlier, is this file sensitive, should this action wait for a person? Those decisions have to be cheap enough to make on every event, and honest enough about their uncertainty to drive a threshold, so we set out to build a small model that answers typed questions with calibrated probabilities. When the first general-purpose decision models appeared, they showed how broadly the same approach applies, and we widened Cox from security events to classification in general. Bringing it back into the sandbox, fine-tuned on sandbox event data, is on our roadmap.

What it does

You give Cox a text and a set of typed questions: yes/no, pick one of these options, or rate on this scale. It returns a probability for every option of every question, from a single pass through a small model (4 billion parameters, running on one consumer GPU). It can’t answer outside the options you gave, so there is nothing to parse and nothing to go off-script.

The probabilities are calibrated: answers given at 80% confidence are right about 80% of the time. That is what makes them usable in code. Automate above a threshold, send the rest to a person, weigh asymmetric costs, or notice when the text supports two answers.

How it works

  • Read, don’t write. We start from an open language model and train it to answer by reading. The text and all the questions go in together, and each answer is read from the model’s internal state at the end of its question. No text is generated.
  • Each question is answered as if asked alone. Asking ten questions gives the same answers as asking them one at a time, but in a single call.
  • Map the question space first. Rather than starting from whichever datasets exist, we asked what kinds of question can be asked about a text: an answer form (yes/no, choice, score) crossed with what it’s about (a property, presence, a relation, compliance with a rule, selection, grading, counting, change, attribution, consistency). That grid shows which skills are untrained and becomes the data plan.
  • Truth by construction. Where no dataset exists, code decides the right answer first, an open model writes natural text around that decision, and a second model checks it blind. The labels are correct by design.

How we measure

A fixed suite of 20 real-world tasks is never trained on: their texts are fingerprinted and refused if they ever appear in training data. One half of the suite uses question shapes the model has trained on, in new settings. The other half uses question shapes it has never seen. For each task we report accuracy, skill (improvement over always guessing the most common answer) and calibration error.

We also check the test labels. Where a model and the dataset disagree, a person reviews the item, and corrections are published as a separate errata list with a fixed rule for what changes. The original data is never edited. Where a label records a fact, such as the star rating a reviewer actually gave, it stays, even when keeping it lowers our own score.

Where it stands

As of October 2026, our best model (an average of several 4B training runs) reaches:

  • skill +0.56 on familiar question shapes in new settings, and +0.61 on question shapes it never trained on;
  • 90% accuracy on items the text clearly answers;
  • calibration error of about 0.01 on long documents it never saw;
  • the same scores at 4-bit (within 0.005 skill), at roughly 2.3 GB.

Not everything worked. Longer training fitted familiar question shapes harder and lost skill on unseen ones. Several plausible ideas did nothing measurable. We record those results alongside the ones that did work.

What’s next

  • New question families: grading against a rubric, attribution (who said or caused what), record matching, and number comparison.
  • A release build trained only on data that permits commercial use, with open weights.
  • Write-ups of the method and results in Articles.

The Python package for running Cox, coxlm, is open source, with worked examples of real inputs and outputs. The weights aren’t public yet; to hear when they are, subscribe or email [email protected].

Related articles

research, cox, evaluation, benchmarks

Seven decision models on one test suite

Cox, with two readout heads, and six openly available decision models, evaluated on the same 21 tasks with the same corrected labels and the same scoring, and compared within three size classes. On the tasks that none of the models trained on, Cox 4B with its pointer head scores highest among the models of 4B parameters or fewer, a small but statistically reliable lead over Kev 4B. Kev 9B, at more than twice the size, scores higher still.

research, cox, interpretability

Where answers form inside a small decision model

We read the answer of a 4B decision model out of every layer of its backbone, for every kind of question it answers. Answers form in the middle of the network and the top third adds little; a hybrid backbone forms them earlier than a full-attention one; averaged models form them in the same layers as their members; answers that end up wrong settle later, though the model's own confidence already captures most of that; and with its fine-tuning switched off, the base backbone already holds much of the answer, which fine-tuning sharpens in the middle layers.

research, cox, evaluation

Mapping the question space: the task-prime matrix

Before training a decision model, we asked what kinds of question can be asked about a piece of text at all. The answer is a small grid, and it changed how we build data and how we measure.