As reliable as an LLM, only faster: grading technical documents with Jev-like System One models
Can a specialised classifier (Jev, simple-jev, D1) judge which sections of a technical document answer a question as reliably as a large language model (LLM), at much lower latency?
Jev and simple-jev missed 1–6 of the 290 sections that answer a question, against 11 for a 31B LLM. Of the irrelevant sections, sections that should have been graded unrelated, they still kept 11–23%. We call this share noise, and set the limit at 25%: at most one irrelevant section in four may get through. Jev graded a question in 19–26 s, at $0.07–0.08 for a full run of 38 questions.
Liquid AI's D1 missed as few answers (3–4), but it kept almost a third of the irrelevant sections (29–30%). Laya, run zero-shot, kept nearly every section. Jev is not clearly cheaper than the LLM we compared it with, Gemma 4 31B. Google has no paid price for that model, so we estimated its cost from its token counts at third-party hosts' prices: $0.06–0.44 per run. Speed is Jev's selling point, and accuracy was the bar it had to clear.
How section grading works
An engineer asks a question about a technical document, for example “what force must a component withstand?”. Before anything can answer it, a system has to decide, section by section, which parts of the document are relevant. That grading step runs once per question–section pair, so at scale it dominates time and cost. LLMs do it well, but slowly. Jev-like classifiers promise the same decision much faster.
Each pair gets one of four labels. Three of them keep the section in front of the reader; only unrelated hides it (Figure 1).
How we evaluated
Reference labels
We wrote 38 questions about two technical documents, with 170 and 47 sections (about 15,000 words, 30 pages). They range from look-ups and lists to reasoning and “does this apply?” questions. Every question was paired with every section: 8,246 pairs.
A Gemini 3.5 flash-lite judge labelled all pairs. Claude Opus 5.5 then corrected those labels in three rounds of blind review of the pairs where systems disagreed with the judge. Table 1 gives the rubric and the final distribution.
| Label | Rubric | Pairs | Share |
|---|---|---|---|
| direct | The section states all or part of the answer. A scope exclusion that answers a “does this apply?” question counts as direct. | 290 | 3.5% |
| tangential | Same subject as the question, but it does not answer it. | 247 | 3.0% |
| reference | Points to where the answer is (another clause, an annex, another standard) without giving it. | 40 | 0.5% |
| unrelated | Everything else. | 7,669 | 93.0% |
Table 1. Label rubric and the distribution of reference labels over all 8,246 pairs. Every grader got the same four definitions.
Metrics
Not every error costs the same. A dropped answer can mean a failed design, whereas a kept irrelevant section costs only reading time. We therefore score two errors separately:
- Missed answers: a direct pair graded unrelated or left unanswered, out of 290. This is the error that matters. A direct answer graded tangential is still shown, so it does not count as a miss.
- Noise: the share of unrelated sections that a system kept anyway. We set the limit at 25%; if more than about one irrelevant section in four gets through, filtering costs more reading time than it saves. The limit is our judgement.
- Kept per question (out of 217) and time per question: summed request latency divided by 38.
Systems tested
- Jev (typesafe.ai, paid API). It takes a text and a set of multiple-choice questions and returns one label per question.
- simple-jev, two open models behind the same interface: Qwen3.8-27B and Gemma-4-26B-A4B (Featherless.ai, free demo).
- D1 (Liquid AI, d1:free), served through the same API as Jev.
- Gemma 4 31B (Google, free tier), the LLM baseline. It answers in free text, which we parse.
- Laya English 0.42B and multilingual 0.32B, an encoder classifier on an 8 GB laptop GPU. We ran the shipped checkpoints zero-shot. Its model card says accuracy comes from fine-tuning on your own domain, which we did not do.
Each request carries one section and k questions, and returns k labels (Figure 2). We ran every system at k = 38 (all questions in one request) and k = 10, to see whether request size changes the result.
Screening pilot
Before the full runs, we screened local SLMs, LLMs and System One models on 480 pairs (10 questions × 48 sections). Three small, local language models that fit an 8 GB laptop GPU (qwen3.5 9.7B, granite4.2-8b, phi4-mini) each either missed many answers (7–8 of 26) or kept almost everything (73% noise for phi4-mini). Given their poor accuracy and high noise, they were no longer considered for the full run. The LLM, Gemma 4 31B, went on to the full set. All System One models, Jev, simple-jev (on Qwen3.8-27B and Gemma-4-26B), D1 and Laya, ran the full set as well, although Laya failed the pilot and D1 was just over the noise limit.
Results
We did one offline run per system on the full 8,246 pairs. D1 is not deterministic (about one label in nine changes between runs), so its figures are the mean of three runs. Figure 3 plots the two errors against each other.
| System | k | Missed /290 | Noise | Noise, corr. | Kept /q |
|---|---|---|---|---|---|
| Jev | 38 | 4 | 22.8% | 16.4% | 60.7 |
| Jev | 10 | 3 | 22.7% | 16.3% | 60.5 |
| simple-jev Qwen3.8-27B | 38 | 5 | 14.1% | 6.9% | 42.9 |
| simple-jev Qwen3.8-27B | 10 | 1 | 15.6% | 8.6% | 45.9 |
| simple-jev Gemma-4-26B-A4B | 38 | 6 | 10.9% | 4.8% | 35.4 |
| simple-jev Gemma-4-26B-A4B | 10 | 6 | 11.4% | 5.6% | 36.3 |
| D1 (mean of 3 runs) | 38 | 3.7 | 29.5% | – | 73.9 |
| D1 (mean of 3 runs) | 10 | 3.7 | 30.1% | – | 75.2 |
| Gemma 4 31B (LLM) | 38 | 11 | 8.1% | 3.8% | 29.2 |
| Laya English, zero-shot | 38 | 0 | 100% | – | 217.0 |
| Laya multilingual, zero-shot | 38 | 28 | 77% | – | 168.2 |
Table 2. Full-set results. Noise, corr.: blind review of a sample of kept “noise” found 28% of it relevant after all, and this column removes that share; D1 and Laya were not in that review. All misses fall in 8 of the 38 questions. D1's misses ranged 2–5 at k38 and 3–4 at k10 across three runs; its noise varied by under 1 point.
Three patterns stand out. Low noise comes with more misses: the LLM and simple-jev Gemma keep the least, and miss the most (6–11). Request size barely matters for Jev (3 vs 4 misses), but does for simple-jev Qwen (1 at k10, 5 at k38). And D1's scores rank answers slightly worse than Jev's (AUC 0.95 against 0.98, direct vs unrelated), so a stricter cut-off brings D1 under the limit only at the cost of 4–8 misses.
Speed and cost
Speed is Jev's selling point. The other systems ran on free tiers or demos, whose times include queueing, so those comparisons are indicative only (Table 3).
| System | Time /q, k38 | Time /q, k10 | Cost of one full run |
|---|---|---|---|
| Jev | 18.8 s | 26.4 s | $0.065 / $0.077 (measured) |
| D1 | (11.6 s) | (32.3 s) | free API |
| simple-jev Qwen3.8-27B | (31.7 s) | (135.6 s) | free demo |
| simple-jev Gemma-4-26B-A4B | (48.3 s) | (172.3 s) | free demo |
| Gemma 4 31B (LLM) | (816 s) | – | free tier; est. $0.06–0.44 |
| Laya, zero-shot | 4.2–4.5 s | – | local GPU |
Table 3. Summed request latency per question, and cost. Parentheses mark free services. Jev's cost is from its API's token counts. Gemma 4 31B has no paid Google price, so its cost is estimated from its tokens at third-party hosts' prices.
Jev is the fastest system that cleared both bars. On cost, it is not clearly cheaper than Gemma 4 31B: $0.07–0.08 per run against $0.06–0.44, estimated from Gemma's token counts at third-party hosts' prices.
Limitations
- One run per deterministic system; D1 is the mean of two. Small differences in missed answers (for example 3 vs 5) are within run-to-run variation for D1, and may be for others.
- Small set: 38 questions and 290 answering pairs, from two documents. All misses come from 8 questions.
- Latency is unequal: only Jev was measured on a paid, dedicated service.
- Reference labels come from LLMs (a judge, corrected by blind review), not from a human domain expert pass over all 8,246 pairs.
- Laya was not fine-tuned. Its result says nothing about a Laya trained on this domain.
- The 25% noise limit is a judgement, not a measured threshold.
Conclusion
On this task, Jev-like classifiers were as reliable as a 31B LLM, or more so: 1–6 missed answers against 11, with noise under the limit. Jev is the fastest system that passed, at 19–26 s per question, but it keeps the most sections of those that passed (61 per question against 29–46), and it is not clearly cheaper than Gemma 4 31B at third-party prices. D1 is fast and misses little, but keeps too much. Laya cannot do the task zero-shot. Next is a latency comparison on paid, dedicated endpoints, which will show whether Jev's speed advantage holds when nobody queues.