Illustration: three small, fast robots labelled Typesafe AI Jev and Featherless.ai simple-jev race past two large, slow chatbot machines that are still thinking, under the banner “Intelligence… just faster.”
Research1 October 2026@AI_Andy__

As reliable as an LLM, only faster: grading technical documents with Jev-like System One models

An offline evaluation of six graders on 38 questions and 8,246 question–section pairs.

Core question
Can a specialised classifier (Jev, simple-jev, D1) judge which sections of a technical document answer a question as reliably as a large language model (LLM), at much lower latency?
A small glowing robot reading a page, asking: same job, faster brain?
TL;DR

Jev and simple-jev missed 1–6 of the 290 sections that answer a question, against 11 for a 31B LLM. Of the irrelevant sections, sections that should have been graded unrelated, they still kept 11–23%. We call this share noise, and set the limit at 25%: at most one irrelevant section in four may get through. Jev graded a question in 19–26 s, at $0.07–0.08 for a full run of 38 questions.

Liquid AI's D1 missed as few answers (3–4), but it kept almost a third of the irrelevant sections (29–30%). Laya, run zero-shot, kept nearly every section. Jev is not clearly cheaper than the LLM we compared it with, Gemma 4 31B. Google has no paid price for that model, so we estimated its cost from its token counts at third-party hosts' prices: $0.06–0.44 per run. Speed is Jev's selling point, and accuracy was the bar it had to clear.

How section grading works

An engineer asks a question about a technical document, for example “what force must a component withstand?”. Before anything can answer it, a system has to decide, section by section, which parts of the document are relevant. That grading step runs once per question–section pair, so at scale it dominates time and cost. LLMs do it well, but slowly. Jev-like classifiers promise the same decision much faster.

Each pair gets one of four labels. Three of them keep the section in front of the reader; only unrelated hides it (Figure 1).

Question 1 of 38 Sections 217, two documents Grader one label per pair direct tangential reference unrelated shown first shown shown hidden
Figure 1. The grading step. A section labelled unrelated never reaches the reader, so if it held the answer, the answer is lost.

How we evaluated

Reference labels

We wrote 38 questions about two technical documents, with 170 and 47 sections (about 15,000 words, 30 pages). They range from look-ups and lists to reasoning and “does this apply?” questions. Every question was paired with every section: 8,246 pairs.

A Gemini 3.5 flash-lite judge labelled all pairs. Claude Opus 5.5 then corrected those labels in three rounds of blind review of the pairs where systems disagreed with the judge. Table 1 gives the rubric and the final distribution.

LabelRubricPairsShare
directThe section states all or part of the answer. A scope exclusion that answers a “does this apply?” question counts as direct.2903.5%
tangentialSame subject as the question, but it does not answer it.2473.0%
referencePoints to where the answer is (another clause, an annex, another standard) without giving it.400.5%
unrelatedEverything else.7,66993.0%

Table 1. Label rubric and the distribution of reference labels over all 8,246 pairs. Every grader got the same four definitions.

Metrics

Not every error costs the same. A dropped answer can mean a failed design, whereas a kept irrelevant section costs only reading time. We therefore score two errors separately:

Systems tested

Each request carries one section and k questions, and returns k labels (Figure 2). We ran every system at k = 38 (all questions in one request) and k = 10, to see whether request size changes the result.

ONE REQUEST Section text one of 217 + k questions each with the four labels k labels k = 38 or k = 10
Figure 2. Request shape. Jev, simple-jev, D1 and Laya received the same request. The LLM received the same content as a prompt.

Screening pilot

Before the full runs, we screened local SLMs, LLMs and System One models on 480 pairs (10 questions × 48 sections). Three small, local language models that fit an 8 GB laptop GPU (qwen3.5 9.7B, granite4.2-8b, phi4-mini) each either missed many answers (7–8 of 26) or kept almost everything (73% noise for phi4-mini). Given their poor accuracy and high noise, they were no longer considered for the full run. The LLM, Gemma 4 31B, went on to the full set. All System One models, Jev, simple-jev (on Qwen3.8-27B and Gemma-4-26B), D1 and Laya, ran the full set as well, although Laya failed the pilot and D1 was just over the noise limit.

Results

We did one offline run per system on the full 8,246 pairs. D1 is not deterministic (about one label in nine changes between runs), so its figures are the mean of three runs. Figure 3 plots the two errors against each other.

over the noise limit 0%5%10%15%20%25%30%35% 024681012 Noise: share of irrelevant sections kept Missed answers (of 290) Gemma 4 31B (LLM) simple-jev Gemma simple-jev Qwen simple-jev Qwen Jev D1 k = 10 k = 38
Figure 3. Missed answers against noise on the full set; lower left is better. Jev and simple-jev missed 1–6 answers and stayed under the limit. The LLM kept the least noise but missed the most. D1 missed few but sits just over the limit. Laya (77–100% noise) is off the chart.
SystemkMissed /290NoiseNoise, corr.Kept /q
Jev38422.8%16.4%60.7
Jev10322.7%16.3%60.5
simple-jev Qwen3.8-27B38514.1%6.9%42.9
simple-jev Qwen3.8-27B10115.6%8.6%45.9
simple-jev Gemma-4-26B-A4B38610.9%4.8%35.4
simple-jev Gemma-4-26B-A4B10611.4%5.6%36.3
D1 (mean of 3 runs)383.729.5%–73.9
D1 (mean of 3 runs)103.730.1%–75.2
Gemma 4 31B (LLM)38118.1%3.8%29.2
Laya English, zero-shot380100%–217.0
Laya multilingual, zero-shot382877%–168.2

Table 2. Full-set results. Noise, corr.: blind review of a sample of kept “noise” found 28% of it relevant after all, and this column removes that share; D1 and Laya were not in that review. All misses fall in 8 of the 38 questions. D1's misses ranged 2–5 at k38 and 3–4 at k10 across three runs; its noise varied by under 1 point.

Three patterns stand out. Low noise comes with more misses: the LLM and simple-jev Gemma keep the least, and miss the most (6–11). Request size barely matters for Jev (3 vs 4 misses), but does for simple-jev Qwen (1 at k10, 5 at k38). And D1's scores rank answers slightly worse than Jev's (AUC 0.95 against 0.98, direct vs unrelated), so a stricter cut-off brings D1 under the limit only at the cost of 4–8 misses.

Speed and cost

Speed is Jev's selling point. The other systems ran on free tiers or demos, whose times include queueing, so those comparisons are indicative only (Table 3).

SystemTime /q, k38Time /q, k10Cost of one full run
Jev18.8 s26.4 s$0.065 / $0.077 (measured)
D1(11.6 s)(32.3 s)free API
simple-jev Qwen3.8-27B(31.7 s)(135.6 s)free demo
simple-jev Gemma-4-26B-A4B(48.3 s)(172.3 s)free demo
Gemma 4 31B (LLM)(816 s)–free tier; est. $0.06–0.44
Laya, zero-shot4.2–4.5 s–local GPU

Table 3. Summed request latency per question, and cost. Parentheses mark free services. Jev's cost is from its API's token counts. Gemma 4 31B has no paid Google price, so its cost is estimated from its tokens at third-party hosts' prices.

Jev is the fastest system that cleared both bars. On cost, it is not clearly cheaper than Gemma 4 31B: $0.07–0.08 per run against $0.06–0.44, estimated from Gemma's token counts at third-party hosts' prices.

Limitations

Conclusion

On this task, Jev-like classifiers were as reliable as a 31B LLM, or more so: 1–6 missed answers against 11, with noise under the limit. Jev is the fastest system that passed, at 19–26 s per question, but it keeps the most sections of those that passed (61 per question against 29–46), and it is not clearly cheaper than Gemma 4 31B at third-party prices. D1 is fast and misses little, but keeps too much. Laya cannot do the task zero-shot. Next is a latency comparison on paid, dedicated endpoints, which will show whether Jev's speed advantage holds when nobody queues.