A decision model reads a text and answers typed questions about it (yes/no, a choice among options, or a rating on a scale), returning a probability for every option. Several small decision models have been released in recent weeks, and we have been developing our own, Cox. This article evaluates seven of them on a single test suite, scored identically, as of 10 October 2026, and compares them within three size classes: a 9B decoder, decoders of 2B to 4B parameters, and encoders under 0.5B.

The headline result: on the seven tasks that none of the models trained on, Cox 4B with its pointer head scores highest among the models of 4B parameters or fewer (skill +0.610), ahead of the next model, Kev 4B (+0.588). The lead is small, but it is not down to chance: when the test items are resampled, Cox stays ahead in 95% of resamples or more. Kev 9B, more than twice the size, scores higher than every smaller model (+0.643).

Because we built both the suite and Cox, the comparison is designed to be checked rather than taken on trust: tasks a model trained on are reported separately, results where other models perform better are stated explicitly, and the limitations are collected in one section at the end.

The models

ClassModelDeveloperSizeBase modelSuite tasks in its training dataEvaluated through
9B decoderKev 9BJared Palmer9BQwen3.5-9B-Base6 of 21its own server
2B to 4B decodersCox 4B, slot headVitund4BQwen3.5-4B-Basenoneour server, bf16
Cox 4B, pointer headVitund4BQwen3.5-4B-Basenoneour server, bf16
Kev 4BJared Palmer4BQwen3.5-4B-Base6 of 21its own server
Strands Decider 2B (v19)AWS Strands Labs2BQwen3.5-2B-Base12 of 21its own server (strands-decider serve)
Jeff 2Bfirelex2BQwen3.5-2B9 of 21its own server
EncodersLayaConvAI Innovations421MModernBERT-largenot publishedits PyPI package (laya 0.3.28)
Julia-1Supersonic Labs144MmmBERT-smallapproximately 4 of 21 (inferred)its own runtime, with the author’s CPU settings

Training overlap is taken from each model’s published data list. Julia-1 does not publish one, so its overlap is inferred from the validation sets it ships; Laya does not publish one either. Kev 9B shares Kev 4B’s published training list. Cox excludes the suite by construction: every suite text is fingerprinted, and the data builder rejects any training item that matches.

Two Cox readouts. Both Cox models use the same base model and read the text and every question in a single pass; they differ in how the answer is turned into probabilities. The slot head compares the model’s answer state with each option encoded separately, on its own. The pointer head compares it with each option as read in context, after the text and the question; each option is read in isolation from the others, so the order in which options are listed cannot affect the answer. The pointer-head model is an average of three training runs of half our usual length, trained with a decaying learning rate on a data mix filtered by licence for research use. Because it answers yes/no questions with the slot readout and other questions with the pointer, a separate temperature was fitted for each readout afterwards, on generated held-out data that no training run used (never on the suite), to correct its confidence.

Method

Suite. 21 public datasets covering intent, topic, tone, entailment, similarity, spam, clause type and slot filling, among others: approximately 24,000 decisions in total. Ten tasks use question formats Cox was trained on, applied to new data (Tier A); eleven use formats it was never trained on (Tier B). The tiers are defined relative to Cox’s training only.

Metrics. For each task we report:

  • accuracy, the proportion of decisions whose most probable answer is correct;
  • skill, the improvement over always giving the most common answer, (accuracy - majority) / (1 - majority): 0 means no improvement, 1 is perfect, and negative values are worse than the baseline;
  • calibration error (ECE), the average gap between stated confidence and observed accuracy. An ECE of 0.05 means, for example, that answers given at 80% confidence are correct roughly 75 to 85% of the time.

Averages are taken over tasks, so each task carries equal weight regardless of size.

Labels. Public datasets contain labelling errors. Items where a strong model disputed the dataset label were reviewed by hand, and corrections were published as a separate errata list under a fixed rule; the original data is not modified. Across these 21 tasks, 148 labels were corrected and 46 items removed, and one dataset (emotion) was withdrawn from the suite. Labels that record a fact, such as the star rating a reviewer gave, were retained even where the text suggests otherwise. The corrections raise all models’ scores by similar amounts and change no ranking; all results below use the corrected labels.

Uncertainty. Every model’s predictions are stored per item and scored by the same code. The 95% intervals come from a paired bootstrap over items: each resample is applied to all models at once, so the intervals on differences between models are directly comparable.

Results on tasks no model trained on

Seven tasks are absent from the training data of every model with a published data list: ATIS slot filling, irony, LEDGAR (100 clause types), SciFact, STS-B similarity, subjectivity and WikiQA.

2B to 4B decoders

ModelSizeSkill (95% interval)AccuracyCalibration error
Cox 4B, pointer head4B+0.610 (+0.588 to +0.630)76.9%0.072
Cox 4B, slot head4B+0.595 (+0.573 to +0.617)74.9%0.059
Kev 4B4B+0.588 (+0.568 to +0.608)76.3%0.071
Strands Decider 2B2B+0.482 (+0.457 to +0.505)70.3%0.155
Jeff 2B2B+0.414 (+0.387 to +0.439)64.8%0.111

Cox 4B with the pointer head leads its class on skill, 0.022 ahead of Kev 4B, and is also the most accurate in it (76.9%, against 76.3% for Kev). The lead is small but reliable: when the test items are resampled, the gap falls between +0.003 and +0.042 in 95% of resamples, so it is consistently above zero. The Cox slot head scores 0.007 above Kev, which is within the range chance could produce, so those two are level on skill; the slot head is also the best calibrated model in this class (0.059, against 0.071 for Kev and 0.072 for the pointer head). Because all three share a base model, this is the most direct comparison in the set: different training methods applied to the same 4B starting point.

The two 2B models trail the 4B models by 0.11 to 0.18 in skill, well outside the intervals. Model size accounts for part of this: in a matched experiment of our own, a 2B and a 4B model trained identically scored similarly on their training distribution but 0.12 apart on unseen tasks.

9B decoder

ModelSizeSkill (95% interval)AccuracyCalibration error
Kev 9B9B+0.643 (+0.621 to +0.664)79.0%0.049

Kev 9B scores higher than every model in the smaller class: 0.033 ahead of the Cox pointer head (the gap falls between +0.015 and +0.054 in 95% of resamples), more accurate, and the best calibrated model in this article. It is Kev 4B’s recipe on a base model more than twice the size, and on these tasks the step from 4B to 9B adds 0.055 in skill.

Encoders

ModelSizeSkill (95% interval)AccuracyCalibration error
Laya421M+0.144 (+0.111 to +0.174)57.4%0.242
Julia-1144M-0.416 (-0.479 to -0.356)35.8%0.488

Laya and Julia-1 are considerably smaller encoders (421M and 144M parameters) and represent a different trade-off. Laya exceeds the most-common-answer baseline on six of the seven tasks, but its result on ATIS slot filling (-1.09) lowers its average substantially. Julia-1 falls below the baseline overall: it assigns “yes” a probability above 0.8 on 82% of the suite’s yes/no questions, while “yes” is correct on 45% of them.

Results on each model’s own untrained tasks

Seven tasks form a narrow basis for comparison. The following table scores each model on every task outside its own training data, alongside both Cox 4B models on the same tasks (skill / accuracy).

ModelTasksThe modelCox 4B, pointer headCox 4B, slot head
Kev 9B15+0.652 / 80.7%+0.626 / 79.2%+0.624 / 78.4%
Kev 4B15+0.615 / 78.7%+0.626 / 79.2%+0.624 / 78.4%
Strands Decider 2B9+0.536 / 71.8%+0.660 / 78.8%+0.636 / 76.2%
Jeff 2B12+0.432 / 62.5%+0.685 / 79.4%+0.661 / 77.0%
Laya21 (overlap unknown)+0.390 / 64.8%+0.646 / 78.8%+0.636 / 77.5%
Julia-117-0.238 / 38.5%+0.639 / 79.8%+0.635 / 79.1%

Across all 21 tasks, including the six it trained on, Kev 4B is level with the Cox pointer head (skill +0.648 against +0.646; accuracy 79.3% against 78.8%) and ahead of the slot head (+0.636, 77.5%); Kev 9B scores +0.683 (81.3%). This is why the comparisons above exclude each model’s training tasks.

Where other models lead

  • A larger model. Kev 9B, at more than twice the size, leads every smaller model on skill, accuracy and calibration on the untrained tasks.
  • Accuracy. Kev 4B is slightly more accurate than the Cox slot head on every task set in this article, and than the pointer head across all 21 tasks, which include six it trained on. On the untrained task sets the pointer head is the more accurate.
  • Many-option classification. On LEDGAR, a 100-way clause-type task, Kev 4B (+0.59) and Strands Decider (+0.55) are well ahead of the Cox slot head (+0.35). Both read options in context, as the Cox pointer head does, which reaches +0.58.
  • Speed. Strands Decider, at half the size, responds roughly three times faster (see the JevBench results below).

JevBench

JevBench is an independent benchmark for decision models maintained by Benchmark Heaven. We ran its public v1 set (231 original scenarios, none drawn from public datasets; the version pinned by the Strands Decider repository) without modification:

ModelCorrect (of 231)Calibration errorLatency, median / 95th percentile (RTX 4090)
Cox 4B, slot head168 (72.7%)0.10773 / 387 ms
Cox 4B, pointer head170 (73.6%)0.12684 / 262 ms
Strands Decider 2B168 (72.7%)0.04821 / 132 ms

The three models are close on accuracy: the Cox pointer head answers 170 scenarios correctly, the Cox slot head and Strands Decider 168 each, differences within noise on 231 questions. Strands Decider is better calibrated on this set, and our run reproduces its developers’ published figure. The Cox heads’ temperatures were fitted on our own generated data, and they carry over less well to these scenarios than to the suite.

Jev, TypeSafe AI’s model and the namesake of this class of model, is not openly available, and we have not evaluated it. On Benchmark Heaven’s current leaderboard (JevBench v1.6.0, scored 5 October 2026), Jev 1.13.0 ranks second of 92 systems on a headline score that weights accuracy on public and sealed sets, calibration, speed and cost equally. That version differs in both questions and scoring from the v1 set above, so the two sets of figures are not comparable. Cox does not yet appear on the leaderboard.

A real-time task: ViZDoom

Decision models are fast enough to act once per frame in a game. In ViZDoom’s defend the center scenario (fixed seed, 20 episodes, 26 rounds of ammunition), each model receives the same text description of the frame and the same two questions at every step: which button to press, and whether the player is in danger.

PlayerScore per episode
Random buttons0.4
Hand-written rule-based bot6.6
Kev 4B3.1
Jeff 2B-1.0 (never fires)
Cox 4B, without game training0.7
Cox 4B, after a short fine-tune on a planner’s decisions6.7
Same, with a fuller frame description (all monsters and the previous frame)22.1
The planner it learned from (uses look-ahead and privileged information)25.0
Cox 4B after the fine-tune, with the fuller frame description: one full episode, recorded on an RTX 4090 (score 20). Below the frame are the model's probabilities for each action at that step; the chosen action is highlighted, and the time each decision took is shown above them. Open the video on its own.

The score is one point per kill, minus one for the player’s death, which ends every episode; with 26 rounds the maximum is 25. The standard frame description omits three of the five monsters, and the improvement to 22 comes from supplying that information and training the model to use it. The fine-tune changed the model’s suite scores by about 0.01. In the recording above, each decision takes a median of 66 ms, about 15 decisions per second including the game itself. These runs used an earlier Cox 4B (slot head), not the pointer-head model in the tables above.

Limitations

  • Single runs. Each external model was evaluated once, at the version listed, through its own server or runtime. Each Cox 4B model is an average of several training runs; the pointer head’s runs used half our usual training length.
  • Our suite. We built the suite, and Cox was designed for this kind of task. Tier B is unseen for Cox; for the other models it is unseen only to the extent their training data allows.
  • Unpublished training data. Laya’s and Julia-1’s training data are not published, so their results on “untrained” tasks cannot be verified.
  • Reproducibility. Cox 4B is a research build: some of its training data does not permit commercial use, and its weights are not public. Its results therefore cannot yet be reproduced independently; those of the other six models can.
  • Calibration as shipped. Calibration depends on the temperatures each model ships with; every model was used as released.

Next steps

We plan an open-weights release of Cox trained only on commercially licensed data, and will rerun these comparisons on it. We are also adding Strands Decider v21 and newer models released since this article first appeared. We will update this article, with a new date, when any of the seven models changes.

Appendix: every task, every model

Skill per task on the corrected labels. † marks a task in the model’s training data (from its published data list; inferred for Julia-1). Laya’s training data is unknown, so none of its entries are marked.

TierTaskKev 9BCox pointerCox slotKev 4BDecider 2BJeff 2BLayaJulia-1
AAG News topic+0.88†+0.85+0.78+0.86†+0.88†+0.80+0.91+0.80†
AANLI round 3+0.43†+0.39+0.38+0.31†+0.19†+0.17†+0.10+0.01†
ABanking77 intent+0.86†+0.70+0.66+0.85†+0.82†+0.15+0.37+0.81†
ACivil Comments: identity attack+0.86+0.73+0.80+0.79+0.77†+0.61†+0.28+0.01
ACivil Comments: threat+0.74+0.60+0.72+0.65+0.27†+0.52†+0.93+0.01
ACivil Comments: toxic+0.26+0.32+0.32+0.44+0.56†+0.52†+0.61-0.29
AContractNLI+0.70+0.68+0.67+0.62+0.54†+0.76†+0.34+0.03
AMASSIVE intent+0.77+0.77+0.73+0.75+0.75+0.35†+0.36+0.43†
ASciFact+0.72+0.63+0.69+0.61+0.05+0.40+0.29-0.18
AYelp stars+0.61†+0.57+0.57+0.62†+0.58†+0.42+0.31+0.11
BATIS slots+0.79+0.77+0.79+0.81+0.78+0.66-1.09-1.37
BBoolQ+0.83†+0.78+0.77+0.80†+0.75†+0.76†+0.68-0.02
BIrony+0.52+0.50+0.46+0.34+0.28+0.21+0.39-0.48
BLEDGAR clause type+0.60+0.58+0.35+0.59+0.55+0.12+0.28+0.27
BPAWS paraphrase+0.58+0.62+0.60+0.58+0.81†+0.86†+0.77-0.01
BPubMedQA+0.47+0.46+0.42+0.42+0.41†+0.30†-0.02-0.88
BSMS spam+0.90+0.94+0.93+0.85+0.89†+0.14+0.73-0.06
BSTS-B similarity+0.38+0.33+0.40+0.35+0.35+0.28+0.02+0.03
BSubjectivity+0.88+0.85+0.86+0.87+0.82+0.77+0.63-0.20
BTREC question type+0.96†+0.90+0.82+0.96†+0.70+0.78+0.81-0.04
BWikiQA+0.61+0.61+0.62+0.55+0.56+0.46+0.49-0.98