Tech article

Decision models that pick options, and the four scoreboards that rank them

Published
Written by
HDATF

A decision model receives the options in advance and returns a probability for each one. Five posts published between 19 September and 8 October 2026 covered the architectures, the public rankings and the training tools. This article reads them as one topic and marks what each source can and cannot support.

A business record and two options feed into a large cube. The cube sends one result to four separate boards, and each board has its own scale with differently spaced marks, so the same result lands at a different height on each one.

The question this article asks

Most AI products ask a language model to write an answer and then read that answer back. A decision model works the other way around. It receives the material, the question and the full list of options up front, and it returns a probability for each option.

Between 19 September and 8 October 2026, five posts covered this one topic from different angles: how the scoring is built, how the public boards rank it, and what it now takes to train one. Read separately they look like five small announcements.

Read together they answer one question: if a product hands its small judgements to a decision model, how would anyone tell whether that model is any good? The short answer is that four public measurements exist and they disagree, because each one measures something different.

Four terms to know first

Decision model
A decision model receives the material to judge, a question fixed in advance, and the list of options. It returns a probability for each option and stops there.
Forward pass
A forward pass is one trip of the input through the model. Writing a sentence takes one pass per token. Scoring a fixed list of options can be done in a single pass.
Calibration
Calibration asks whether a stated confidence matches the observed hit rate. A model that says 0.9 on a hundred cases and gets about ninety of them right is well calibrated. Confidence on its own stays a number the model produced.
LoRA
LoRA adds a small set of trainable weights beside a frozen base model. Training touches only the added weights, so the memory a run needs drops a long way.

One refund enquiry, from input to output

Take a single customer enquiry arriving at a support desk. A language model would be asked to write a reply and a label. A decision model is asked something narrower.

State
The enquiry text, the order record it refers to, and the account history. Cloudflare's published code calls this the record and encodes it once.
Questions
Is this a refund request? Does it fall inside the return window? Does it need a human? Each question is written once and reused.
Options
The published code uses three question types. A yes-or-no question, a choice among named options, and a score read off an ordered scale.
Output
A probability per option. In the published code a yes-or-no question comes back as the probability of true rounded to four places, a choice comes back as the chosen key, and a score comes back as the expected value over the scale.
The record is encoded once. Each question is scored against its own options in the same forward pass, and each question gets its own probabilities.

The output carries no sentence to parse and no instruction to obey. A probability of 0.93 on the refund question tells an operator how the model scored that option. Acting on it stays a separate decision the product makes.

Three structures that score the options

Public model cards and published code describe three ways to get a probability per option. The labels below come from the sources themselves.

Reading the option letter
The model is shown the options as lettered choices and the scores it assigns to those letters are read directly, before any sentence is written. Quyet-1.0-Large describes its output this way, as option-letter logits.
A head built for the task
A separate scoring layer sits on top of the base model and produces the probabilities itself. Cloudflare's Clef calls its layer the joint schema head.
Comparing two vectors
The material and each option are turned into vectors and the closest option wins. Laya's card names ModernBERT-large as the encoder it uses.

The published Clef head is the one of the three whose code is readable in full, so it is worth following once.

  1. The record is encoded once. The spans of hidden states belonging to each question and each option are pooled into one vector apiece.
  2. A lexical vector is read off the language model's own output embedding weights, so the words of the option still carry weight.
  3. An evidence layer attends across the record, and a joint field transformer decodes the question and option fields together.
  4. The final score adds three parts: the lexical prior, the joint cosine score, and a residual term. A softmax runs per question, so every question is answered in the same forward pass.

Two limits sit next to that. Clef's encode_record default in the published code is 16384 tokens, while the hosted service advertises a larger window, so the served limit and the function default are separate numbers. And an interface that accepts the same request shape as another product shares the shape, while the weights behind it stay the vendor's own.

Four scoreboards that measure the same model

Four public measurements of decision models exist. Each one was read on 8 October 2026. They produce different orderings because each defines its score differently.

The same result measured by four scales. The heights shown here are drawn for illustration, and they stay separate because the four measurements define their scores differently.
The four public measurements, as each one describes itself
MeasurementWhat it reportsWhy its number travels poorly
JevBenchCapability Score and Composite Score, on a sealed poolIts method moved to version 1.6.1, and cost and latency gates now sit beside the score. A model can hold a high capability number and still fall outside a gate.
S1BenchMacro accuracy across six public suites, with a delta against published figuresIts arms run on hosted, local and CPU setups with different context and truncation settings. Macro accuracy averages the suites rather than pooling the samples.
Decision IndexA wider suite at version 0.2.1, run by the vendorThe vendor ran it and reports it. A shortlist of ten tasks in a blog post covers less than the full index.
JevArenaA separate arena figure quoted on model cardsIt sits on a different scale from the index score on the same card, so the two cannot be added or compared.
  • Re-running changes the number on its own. On S1Bench the same published model scored 0.7751 against its own published 0.7682 over the same six suites, a gap of 0.69 percentage points that came from re-measurement rather than from any change to the model.
  • A snapshot is a snapshot. Laya's 62.5 percent matches this S1Bench run on this date, and the figure belongs to that run rather than to the model in general.
  • September rankings have already moved. One of the five posts quoted a September board where one model led at 75.3 and another followed at 74.6. The board read on 8 October orders those rows differently, and the same capability number now sits outside a cost gate.
  • Speed comparisons in these posts mix serving setups. One chart places hosted API calls, local GPU runs and CPU runs on the same axis, so a ratio read off it describes those deployments rather than the models' intrinsic speed.
  • A popularity list ranks demand. One post showed an OpenRouter screenshot ordering eight decision entries. The screenshot counts requests, so it reports what people called most on that day.

What makes direct training possible

The fifth source is a training tool rather than a model. Unsloth published a decision model path that loads a Clef-style backbone, attaches the joint schema head, and trains with LoRA.

Reading the published loader settles what actually gets trained. When a run loads without full fine-tuning, the base weights are set so they stop receiving updates. The configured path then puts LoRA adapters over that frozen base and trains those adapters together with the head, which is held in 32-bit. A head-only run exists as a separate option.

The same tool exposes a calibration step that reports held-out accuracy, expected calibration error and loss. That step is what turns a trained head into something a team can argue about with numbers.

The base stays frozen while the added weights and the head train. Calibration reports held-out accuracy, and the gate holds anything below the threshold for a person.

The memory and accuracy headlines attached to that announcement are reported by the tool's authors. The figures were taken as the authors' own claim, because the page carrying the full run was inaccessible when this article was written.

What to be careful about

  • A classification carries no permission to act. A model returning the refund option scored a field. Transfers, deletions and deployments need their own approval and their own check.
  • Confidence and observed accuracy are two different numbers. A stated 0.93 becomes meaningful once a held-out set shows how often 0.93 was right.
  • Scores from the four measurements stay separate. Adding them, averaging them or merging them into one ranking produces a number that none of the four defined.
  • Decontamination is a claim made by whoever ran the evaluation. Treat it as the evaluator's statement until an independent run reproduces it.
  • A catalogue that omits a name proves only what that catalogue listed. The official documentation states that the default model list returns text models, so a name missing from it still needs a direct check.

A sequence you can try now

  1. Pick one judgement your product already makes by hand, and write its question and its options down in one line each.
  2. Collect a few hundred past cases with the answer a person gave, and hold a portion back so it never reaches training.
  3. Measure a plain baseline first. Keyword rules and an existing language model prompt both give a number to beat.
  4. Run the held-out set through a decision model and read accuracy and calibration error side by side, in Korean on Korean material.
  5. Decide three thresholds before shipping: above which the result is used, below which it is held, and when it goes to a person.

The last step is the one that keeps the rest safe. A held result costs a little time, and a result used at the wrong threshold costs whatever the action behind it costs.

What is still unconfirmed

  • The full training run behind the tool's memory and accuracy headline. The page carrying it stayed inaccessible, so the figures are recorded as the authors' report.
  • Whether a single forward pass handles a long list of questions at the cost quoted for one question. The published median latency describes one question on one GPU.
  • The evaluation sets and the aggregation code behind the sealed pool. The board publishes scores while the pool stays closed.
  • Several product names in the popularity screenshot. They were left as screenshot entries, with no claim either way about whether a model stands behind each name.
  • Korean-language accuracy for any of these models. None of the four measurements reports a Korean suite, so a team would have to measure it on its own material.

Sources