Most models that classify text are trained on whatever labelled datasets happen to exist: sentiment here, topics there, a few hundred tasks collected from benchmarks. The result is a model shaped by the accidents of what people chose to label. It can be excellent at the questions those datasets ask, and you have little idea which questions it has never been taught.

When we started Cox, a small model that answers typed questions about text with calibrated probabilities, we asked a different question first: given a piece of text, what kinds of question can be asked about it at all?

Two short lists

Every question Cox answers is pinned down by two things, each drawn from a short list.

The answer form. A decision comes out in one of three shapes, which we call primes:

  • yes / no: a probability that the answer is yes;
  • choice: a probability for each of a set of named options;
  • score: a probability for each level of an ordered scale, so the answer can sit between two levels.

Everything else a caller might want, such as several yes/no features at once or a score per aspect, is built from these three.

What the question is about. We call this the target, and ten of them cover what we have seen so far: a property of the text (topic, tone, urgency); the presence of something in it (a date, an amount); a relation between two texts; compliance with a rule carried in the question; selection of the best item from a set; grading against a rubric; counting; change between two versions; attribution (who said or did what); and consistency (whether the parts of a text cohere).

Cross the two lists and you get a grid. Each cell is a skill. Hover over a cell to see an example of the question it stands for.

what it's about ↓ answer form → Yes / no a probability of yes Choice a probability for each option Score a probability for each level of an ordered scale
Property an attribute of the text: its topic, tone, urgency, intent
Presence whether something is in the text: a date, an amount, a name, a claim
Relation how two texts relate: support, contradiction, paraphrase, similarity
Compliance whether the text satisfies a rule carried in the question
Selection which item of a set best meets a criterion
Grading quality against a rubric or a reference
Counting how many of something, in buckets
Change what differs between two versions, and how much
Attribution who said or did what, or what caused what
Consistency whether the parts of a text cohere

Hover over, focus or tap a cell to see an example. N/A cells are not gaps: each is a remapping of a cell that exists.

What the grid showed us

The first time we placed our training data in the grid, the picture was lopsided. Almost everything sat in the top row: properties of a text, asked as yes/no or as a choice. Of the 176 instruction-style tasks in our mix, 135 were binary. Whole rows were empty. Nothing taught the model to grade against a rubric, to count, to compare two versions, to say who agreed to what, or to pick the best item from a list. Graded relations (how similar are these two texts?) had no training data at all.

None of that was visible from a list of datasets. It was obvious from the grid.

The target decides where the truth comes from

The second thing the grid gave us was a plan. Each target implies where a correct label can come from:

  • For presence, counting, change, compliance and consistency, code can decide the right answer first: insert a date or don’t, put four action items in the notes, edit one clause, apply the rule, break one sentence. A language model then writes natural text around that decision, and a second model checks it blind. The label is right by construction, not by someone’s opinion.
  • For property and grading, the label is a judgement, so it needs people or a strong teacher model, and it carries their uncertainty.
  • Relations sit in between: some are structural, some are judgements.

So the empty cells became a generation agenda. We built families of examples for counting, change, compliance, selection, grading, attribution, record matching and number comparison, each targeting cells the existing data left empty.

Four cells that are not gaps

Four cells are marked N/A: presence, compliance, selection and attribution, each asked as a score. At first they look like gaps. They are not questions of their own.

“How present is a date, on a 0-3 scale?” is the yes/no presence probability mapped onto the caller’s own scale. “Rate this claim’s compliance from 1 to 5” either rescales the yes/no answer or hides several rules inside one number. Because the probabilities are calibrated (an answer given at 80% is right about 80% of the time), callers can draw those lines themselves, and the model needs nothing new to learn.

Measuring by cell

The grid also changed how we measure. “How good is the model?” has no single answer. “Can it count? Can it grade against a rubric it has never seen? Can it tell which of two people agreed?” do. We score every question in our evaluation under its cell, keep real-world test data apart from generated data, and look at the grid rather than a single average. A strong average can hide an empty row.

Here is what that looks like for our current research build (9 October 2026), on the generated families built to fill the empty cells. New texts are unseen examples from the same domains the model trained on; new domains are whole domains held out of training (a counting family trained on lists of fruit and cities, tested on tools and rivers, for instance). Skill is 0 for always giving the most common answer and 1 for always being right; calibration error is the gap between stated confidence and observed accuracy.

Family (target)New texts: skillNew domains: skillNew domains: calibration error
Change between versions+1.00+0.990.006
Counting+1.00+0.990.006
Compliance with a rule+0.98+0.740.061
Selection of the best item+0.97+0.790.094
Graded levels (urgency, severity, formality)+0.97+0.930.070
Relations between two texts+0.94+0.960.009
Number comparison+0.80+0.880.059
Record matching+0.59+0.540.098

Two things stand out. Where the label is decided by code, the model learns the skill almost perfectly and carries it to new domains (counting, change, relations). Where a skill depends on careful reading of every field (record matching) or on applying an unfamiliar rule in an unfamiliar domain (compliance, selection), new domains are markedly harder, and those are the rows we are now generating more data for. These are our own generated tests, written in the format the model was trained on; real-world datasets are scored separately, in our comparison of decision models.

What the grid is not

It is not a taxonomy of everything language models do. Cox reads and decides; it doesn’t write, so generation and free-text extraction are outside it by design. Two further axes refine a cell: the shape of the input (one text, a pair, a list, a record, a conversation, code), and the source of truth above. And ten targets are where we are today, not a claim that the list is complete.

But if you are building a model that makes decisions about text, or a benchmark that tests one, placing each question in the grid is a cheap exercise. It tells you what you are measuring, and, more usefully, what you are not.