Jev Made "System One" AI Famous. In Thai, Confidence Is the Part to Check.
A System One model decides alone when its confidence is high. We tested this in Thai, on AI an organisation can own. Out of the box, the confidence was often wrong on social-media posts. One simple adjustment, made with a small set of labelled examples, made the confidence honest for every model we tested. Check the confidence on your own data before you let a model decide alone.
Infozense · CIO · Head of Data · ~18 minutes · Tested September 2026
Who this is for
This article is for CIOs, heads of data and IT leads in organisations that need AI to work in Thai, or in any language other than English. We use Thai as the test case. You may have seen the attention around Jev and "System One" models. You want the same benefit, and you want to know whether you can have it on AI your organisation owns or controls. This article reports what we measured when we tried the idea in Thai.
Jev makes the right point: most AI work is deciding, not writing
In September 2026 TypeSafe AI released a model called Jev. It is not designed for conversation. You send a question with a fixed list of answers, and Jev picks one answer and returns a probability for every option. TypeSafe says end-to-end response time is 70 to 500 milliseconds. TypeSafe calls this kind of model "System One", after the fast, intuitive thinking Daniel Kahneman describes in Thinking, Fast and Slow.
The idea behind it is right. Look at what organisations actually want from AI. Which department should handle this request? Is this complaint urgent? What kind of document is this? Each is a decision with a few possible answers. None of them needs a model that can write an essay.
What makes this work is the probability. TypeSafe says every Jev answer comes with calibrated probabilities and confidence scores. (Calibrated means that when the model says 90%, it is right about 90% of the time.) Reliable confidence lets the fast model answer on its own when the score is high enough, and send the rest to an LLM.
Jev is a managed service. TypeSafe runs it for you, and you pay for what you use: TypeSafe lists a price per million input tokens (a token is a small piece of text used to count usage). That makes it easy to start. Some organisations want more than that: they want to fine-tune the model on their own examples, or keep it on AI they own or control. For them there is another choice, and the rest of this article is about it.
What "System One" and "System Two" mean
In Thinking, Fast and Slow, the psychologist Daniel Kahneman describes two ways people think. He calls them System 1 and System 2.
System 1 is fast and automatic. You use it when you read a familiar word, recognise a face, or answer 2 + 2. It costs almost no effort, and it is right most of the time on familiar things.
System 2 is slow and deliberate. You use it when you multiply 17 by 24, or check a contract clause by clause. It takes effort and attention, so you use it only when you need to.
Most of your day runs on System 1. System 2 steps in when System 1 is unsure, or when something does not fit. That split is what makes people efficient: careful thinking is kept for the cases that need it.
AI can work the same way. An AI "System One" makes routine decisions fast, by picking from a fixed list of answers and saying how sure it is. An AI "System Two" is a slower, more capable model that takes the cases the fast part is unsure about. The rule that links them is confidence: when the fast part is sure, it answers; when it is not, it hands over. An organisation works the same way: a front desk handles routine requests and passes the unusual ones to a specialist.
TypeSafe named its model class after this idea. In this article we build both parts on AI an organisation owns: an encoder as System One, and a compact LLM as System Two.
Two kinds of language model: encoder and decoder. The fast part is an encoder. It reads the whole text in one pass: each word looks at the words before and after it, and keeps its position, so word order still counts. It gives each option a score, but it cannot write. The encoder we used is compact, 1.2 GB, and answers in about 11 milliseconds; that is a design choice, not something every encoder has. The careful part is a decoder, the kind of model behind chat AI and most LLMs. Each word looks only at the words before it, and the model writes by predicting the next word, one word at a time. In our test we asked it for a single word, the option letter, so its 36 milliseconds come from its larger size, 7.6 GB, and its longer input (the prompt carries the instructions and every option), not from writing long answers. The two are separate models. The encoder passes on only its answer and how sure it is, so either part can be replaced by a better one later. It is easy to assume that Jev is an LLM, but it is not a chat model: TypeSafe says it gives up text generation and returns structured answers. TypeSafe has not published how Jev is built. Laya, an open-source model that follows the same approach, is an encoder.
Why the confidence number is the part that matters
Everything in the System One idea hangs on one number: the confidence, the percentage that comes with each answer, such as "smart home 96%" in the first figure. If that confidence is at least 90%, the system uses the answer without checking it again. We use 90% in this article, but it is our choice, not a standard. A higher threshold lets fewer wrong answers through but sends more work to the LLM; the right level depends on what a wrong answer costs you. Whatever level you pick, check that the confidence is honest: for example, whether its 90% really means 90%.
If it does not, the routing goes wrong. A model that says 90% but is right only two times in three lets through one wrong answer in three, and nobody sees them. A model that says 60% when it is nearly always right hands over work it could have done itself.
Most published comparisons of System One models rank them by accuracy. We checked their confidence as well.
What we tested
Does a System One model's confidence hold up in Thai? And can the whole pattern run on AI your organisation owns or controls, without relying on someone else's service?
The short answers. Yes, it runs on AI you own: everything in this article ran on one GPU. But out of the box, no model's confidence held up on social-media posts. On routine requests every model passed the 90% check, though most were still somewhat overconfident. One cheap step handles both: adjust the confidence with a small set of labelled examples, and every model's confidence then matches how often it is right.
We tested six models:
An open-source System One model, with no training by us: Laya, from Convai Innovations, released in September 2026 under the Apache 2.0 licence.
Our own System One: an encoder (mmBERT) trained on 200 labelled examples.
Three LLMs with no training: Typhoon 2.5 4B, which is trained specifically for Thai, and the generic Qwen3 4B and Qwen3 8B. An LLM is a large language model that can read and write text; 4B means 4 billion parameters.
Typhoon 4B fine-tuned on the same 200 examples, as case one describes.
We tested them on a single 24 GB NVIDIA workstation GPU (the GPU is the processing card AI runs on). We used two public Thai datasets:
MASSIVE (Amazon, CC BY 4.0): short everyday requests to a voice assistant in 8 topics, such as "set an alarm for nine am" (alarm) or "can they provide takeaway" (takeaway). We tested on 400 Thai requests at a time.
Wisesight Sentiment (CC0): real Thai social-media posts in 3 sentiment groups. We tested on 450 posts at a time. This set is hard even for specialist models. WangchanBERTa is a Thai model widely used as a reference in Thai research. In its published paper it reached 76.19 micro-F1 after training on all of the roughly 21,600 posts. That work used 4 sentiment groups and a different metric from ours, so the two results cannot be compared directly.
Every result in this article is the mean of 3 random test samples. Accuracy differences were checked with a paired bootstrap over test items; the calibration figures are means without intervals. Their accuracy is in the table near the end. This article is about their confidence.
Out-of-the-box confidence failed on Thai posts
For each model we counted the answers it gave with at least 90% confidence, and how many of those were wrong. On social-media posts, out of the box, every model failed this check: of the answers given with at least 90% confidence, 25% to 47% were wrong, depending on the model. The chart in the next section shows a different number: confident wrong answers as a share of all posts. That share also depends on how often a model's confidence reached 90%. Laya's confidence reached 90% less often than the others', so its confident wrong answers came to 9.7% of all posts. Qwen3 4B was the worst: 45% of all posts got a confident wrong answer.
Training on the task did not prevent it. Our encoder was trained on 200 posts from the same dataset, and it was among the most overconfident: 27% of all posts got a confident wrong answer. Fine-tuning on few examples is known to make models overconfident.
On routine requests every model passed this check: of its answers given with at least 90% confidence, at least 90% were right. Most were still somewhat overconfident (calibration error 0.05 to 0.15), and between 3% and 9% of all requests got a confident wrong answer.
For the LLMs, we read the confidence from the probabilities of the option letters. That is what a developer would compute, not a number the vendor ships. This is one common way to read an LLM's confidence; others, such as using the whole vocabulary or changing the order of the options, can give different results.
One number made the confidence honest
The fix is cheap, and it worked on every model. It is called temperature scaling. You take 150 to 200 examples whose answers you know, and fit one number that softens the model's probabilities until its confidence matches how often it is right. The model and its answers stay the same. Only the confidence changes. The method comes from Guo et al. (2017), On Calibration of Modern Neural Networks, which also found that modern neural networks tend to be overconfident.
After one fit, for all six models, the share of all routine requests that got a confident wrong answer fell to at most 2.2%, and on social-media posts to almost zero. Look at how. On routine requests the models still reached 90% confidence on most requests (41% to 90% of them, depending on the model). On posts, almost no answer reached 90% any more (0% to 2.1% of posts), so almost nothing would be decided automatically. The calibration error, which measures the gap between confidence and accuracy, fell from between 0.23 and 0.47 to between 0.04 and 0.11 on social-media posts.
Take our encoder on routine requests. The fit lowered the share of requests it was at least 90% sure of from 94% to 72%, and raised its accuracy on that share from 92% to 99%. The requests it was no longer sure of went to the LLM instead of slipping through.
After the fit, the models were almost never 90% sure of a social-media post (0% to 2.1% of posts), so at this threshold almost nothing would be decided automatically. That is the honest result for this task with 200 examples: none of the setups we tested is reliable enough to decide alone at 90%. Route those answers to a person, accept the accuracy of the best option (67.2% for the fine-tuned Typhoon 4B), or check accuracy at a lower threshold before ruling automation out.
Accuracy does not change with the fit. What changes is which answers you can let through unchecked.
Calibration holds only for the task and data it was fitted on
The fix has one condition: fit it on the same decision and the same kind of text the model will actually see. Laya, Typhoon and the two Qwen models answered both kinds of text as the very same model, so we could test this. We fitted the temperature on routine requests only, and then used that setting on social-media posts.
It did not carry over well. Calibrated on routine requests, Typhoon 4B was excellent there: only 1.8% of requests got a confident wrong answer. On social-media posts the same setting helped, but not enough: 15.7% of all posts still got a confident wrong answer (25.6% out of the box), and of its answers given with at least 90% confidence only 73% were right. Calibrated on the posts themselves, it gave no confident wrong answers at all.
It fails the other way too. Calibrated on social-media posts and used on routine requests, Typhoon was never 90% sure of anything, so it would have handed every routine request to the next step and thrown away the speed. The Qwen models behaved the same way, and Laya less strongly.
So calibration belongs to the model and the decision together, not to the model alone. Our two datasets differ in task, number of options and difficulty as well as in text, so this test cannot say which of these matters most. The practical rule is the same: fit and check the temperature on the exact decision and data the model will see. Related research on shifted data points the same way: Ovadia et al. (2019), Can You Trust Your Model's Uncertainty?, found that calibration fitted on one kind of data falls short when the data shifts.
Two cases where a System One earns its place
Why pair a System One with an LLM at all? An LLM trained for Thai is accurate on its own: Typhoon 4B answered 91.5% of routine requests correctly with no training (our test; all results are in the table further down). A calibrated System One in front earns its place in two cases, one in each of the next two sections. On high-volume routine requests, it answers most of them itself, in about 11 milliseconds, and leaves the LLM free for the hard ones. And when you have to keep a generic LLM, it brings it level with a model trained for Thai.
Case one: with calibrated confidence, System One answers most routine requests itself
The System One idea does not have to live inside a single model. You can build it from two parts your organisation owns: a fast part that answers when it is sure, and a careful part for everything else.
System Two: a compact LLM, fine-tuned. We fine-tuned Typhoon 4B on 200 examples that already had their answers labelled. The technique was LoRA (a fine-tuning technique that trains only a small set of added parameters and leaves the original model untouched). On our 24 GB workstation GPU, training took 75 to 103 seconds and less than 10 GB of GPU memory. The encoder, trained on the same examples on the same card, took 12 to 20 seconds. Times will differ on other GPUs. The fine-tuned model answered 92.8% of routine requests correctly, and its 67.2% on social-media posts was the highest in our tests.
System One: an encoder in front. An encoder is a compact model that reads text and picks an answer from a list, without writing any text. Fine-tuned on the same 200 examples, the encoder used 1.2 GB of memory and answered in about 11 milliseconds. On its own it was right 89.6% of the time on routine requests and 60.0% on social-media posts. The fine-tuned 4B got 92.8% and 67.2%, with more than six times the memory and about three times the time. There is one rule. If the encoder is at least 90% confident, it answers itself. If not, the request goes on to System Two.
On routine requests the encoder answered 72% of requests itself. When it was that confident, it was right 98.3% to 99.6% of the time. Overall accuracy was 93.2%, with no significant difference from the fine-tuned LLM alone (92.8%). Estimated average response time, from the measured per-model times, fell from about 36 to about 21 milliseconds. Memory goes up instead. Both models stay loaded, so the pair needs 8.8 GB of GPU memory, against 7.6 GB for the fine-tuned LLM alone. What the encoder saves is time, not memory; only an encoder used on its own saves memory.
When text is varied, System One steps aside. On social-media posts the encoder reached 90% confidence on fewer than 1 post in 100. So it sent almost every post on to System Two. Accuracy stayed at 67.2%, but every post now went through both parts, so response time rose from 36.5 milliseconds for the fine-tuned LLM alone to 47.2 milliseconds.
That is a useful signal. The encoder's confidence is calibrated, so if System One is almost never confident on your data, the task is too hard for it with the examples you have. Let the fine-tuned LLM decide on its own.
Case two: when you must keep a generic LLM, a Thai encoder in front brings it level
Why not simply use Typhoon? If you can choose, do: a model trained for Thai on its own is the simpler answer. But many organisations already run a generic LLM such as Qwen for other work, and have tested, approved and secured that one model. Adding a model trained for Thai would mean running a second 4B LLM: about 7.6 GB more GPU memory, and one more model to test, approve and update. An encoder adds about 1.2 GB. Can it close the Thai gap instead?
We tested this with Qwen3 at 4B, the generic model that Typhoon 2.5 is built on, with no extra training. On its own it answered 87.3% of routine requests and 52.1% of social-media posts correctly. That is 4.2 and 10.6 points below Typhoon, and both gaps are outside statistical noise. Same size, same design: what differs is Typhoon's additional training for Thai.
Then we put the encoder from case one in front of it, trained on the same 200 Thai examples. Qwen itself was not trained at all. The encoder does two jobs. When it is at least 90% sure, it answers itself. When it is not, it gives Qwen a shortlist of its most likely options, 3 of the 8 topics or 2 of the 3 sentiments, and Qwen picks from that shortlist.
With the encoder in front, Qwen3 4B reached 91.9% on routine requests and 61.7% on social-media posts. That is level with Typhoon alone, at 91.5% and 62.7%. Neither difference is statistically clear: the 95% confidence intervals are −1.1 to +1.9 points and −3.4 to +1.5 points. Note that the pair used 200 labelled examples and untrained Typhoon none; Typhoon fine-tuned on the same examples reached 67.2% on posts. So the encoder does not make a generic model better than one trained for Thai. It makes it as good, without changing the generic model.
On routine requests the pair was also faster, because the encoder answered most requests itself: about 21 milliseconds on average, against about 36 for Typhoon alone. On social-media posts the shortlist did the work. The 90% rule alone added almost nothing (52.2%), because the encoder was rarely that sure. Every post went through both parts, so the pair was slower there than Typhoon alone: 45 milliseconds against 36. The pair needs 8.7 GB of GPU memory, against 7.6 GB for Typhoon alone.
The practical point is that the encoder does not change Qwen at all. The same model keeps serving the rest of your work. When a new version of Qwen comes out, you can swap it in without retraining the encoder. You only measure again.
Every result in one table
Stage
Option
Routine requests (MASSIVE)
Social-media posts (Wisesight)
Time per request
GPU memory
No extra training
Laya (open-source System One model)
68.9%
51.1%
about 13 ms / 12 ms
1.2 GB
Generic 8B LLM (Qwen3 8B)
87.5%
62.3%
57.6 ms / 38.9 ms
15.3 GB
Generic 4B LLM (Qwen3 4B)
87.3%
52.1%
34.7 ms / 34.5 ms
7.5 GB / 7.6 GB
Thai 4B LLM (Typhoon 2.5)
91.5%
62.7%
35.7 ms / 35.9 ms
7.6 GB
Trained on 200 examples
Encoder (mmBERT)
89.6%
60.0%
10.7 ms / 10.9 ms
1.2 GB
Thai 4B LLM, fine-tuned with LoRA
92.8%
67.2%
36.1 ms / 36.5 ms
7.6 GB
Both together
Encoder in front of the fine-tuned 4B LLM
93.2%
67.2%
21.0 ms / 47.2 ms
8.8 GB
Encoder in front of the generic 4B, untrained (shortlist)
91.9%
61.7%
20.6 ms / 45.2 ms
8.7 GB
Mean of 3 random test samples. Time measured one request at a time on one 24 GB workstation GPU; times will differ on other GPUs. Where a cell has two values, the first is routine requests and the second is social-media posts. GPU memory is the peak memory PyTorch allocated for one request, model weights included; a GPU needs additional memory on top of this for the CUDA runtime and PyTorch's cache, so size hardware with headroom. The encoder ran in 32-bit precision, the LLMs in 16-bit.
Before you let any System One decide alone
Whichever model you choose, open or managed, check it on your own work first:
Collect labelled examples of the decision from your real work, in three separate sets: about 200 to train on if you fine-tune, about 200 to fit the confidence, and a few hundred more to check the result. That is roughly 600 or more in total.
Measure accuracy on the check set.
Check the confidence on the check set. Of the answers given with at least 90% confidence, count how many are wrong.
If too many are wrong, fit one temperature on the second set, never on the training examples. It changes no answers.
Set the threshold on the calibrated number, and check it on the check set. Above it, System One answers. Below it, an LLM or a person decides. Look at two numbers: how many answers stay above the threshold, and how many of those are wrong.
Check again when your data changes, or when you change the model.
Which AI setup fits your work: a decision map
Put together, our results give a short decision map. Every number in it comes from our tests. The map shows the main paths; item 2 below is a cheaper variant of item 5.
If your text is in English, an encoder on its own may be enough. With no training, Laya answered 85.3% of the English routine requests correctly. Encoders trained on about 150 English examples reached 93.2% to 93.9%, with 1.2 GB of GPU memory. We did not test LLMs in English, so this branch compares encoders only.
If about 90% accuracy is enough, or your hardware is tight, an encoder trained on your examples can decide alone. On routine requests ours was right 89.6% of the time, in about 11 milliseconds with 1.2 GB of GPU memory. Or let it answer only when its confidence is 90% or more: it then answered 72% of routine requests, right about 99% of the time, and flagged the other 28% for a person to check.
If your text is in Thai or another language, and you can choose the LLM, use a compact LLM trained for that language. Then fine-tune it on about 200 examples of your organisation's decision. On one workstation GPU this took under two minutes.
If you must keep a generic LLM, put an encoder trained on your language in front of it. In our test this brought Qwen3 4B level with Typhoon, without changing Qwen.
If requests are high in volume and routine, put an encoder in front of the fine-tuned LLM. That encoder is your System One. It answered 72% of routine requests itself and cut the average time from 36 to 21 milliseconds.
If the text is varied, let the fine-tuned LLM decide alone. Calibrated confidence tells you which case you have: if System One is almost never sure on your data, it only adds time.
Keep every part on AI your organisation owns or controls. Every branch runs on a single 24 GB GPU.
Limits of this test
We did not test Jev, and we do not claim to be more accurate than Jev.
Our main results are in Thai only. The pattern should carry over to other languages beyond English, but the numbers will not. Measure in your own language.
Laya's developers do not publish its training data, so it may have seen the MASSIVE requests during training. That would flatter its English score most.
In English we tested encoders only, on routine requests. We measured their accuracy, not their confidence. We did not test LLMs in English.
We did not fine-tune the generic Qwen3 4B, so we cannot say how a fine-tuned generic model would compare with an encoder in front of it. We also did not test Qwen's other abilities: "does not change Qwen" means its weights are not touched.
We tested two public datasets. The texts are short, single messages, not conversations.
Our 200 examples came from those datasets' training sets, not from your data. Your work will give different results, and the only way to know is to measure on your own examples.
We timed one request at a time in plain PyTorch (PyTorch is the standard library for running AI models) on one workstation GPU. Times depend on the GPU: other cards will be faster or slower, and on a tuned production server every option would likely be faster.
Every result is the mean of 3 random samples. Trained models vary from sample to sample, depending on which examples they get, and our significance tests do not fully capture that. So we report close comparisons as no clear difference, never as a win.
For the LLMs, out-of-the-box confidence is our reading of the option-letter probabilities. Their vendors do not ship it as a calibrated number.
We did not test Jev's confidence. TypeSafe says its probabilities are calibrated; check them on your own data, as with any model.
Our test sets were balanced across options. Real work is not, and calibration depends on the mix of cases. Fit and check on a sample of your real work.
What we present is a pattern and a measurement, not a ready-made service.
These results hold for the six models we tested. In a later test, a frozen embedding model with a simple classifier on top was not overconfident out of the box. Check whichever model you use.
The bottom line
Jev is right that most AI work is deciding, not writing, and that a fast model can decide alone when its confidence can be trusted. In Thai, out of the box, the confidence of all six models we tested failed on social-media posts: their confident answers were wrong a quarter to almost half of the time, even for a model trained on the task. One number, fitted on 150 to 200 labelled examples of the same kind of text, made the confidence of all six honest. With honest confidence, a System One in front of an LLM answered 72% of routine requests itself, kept accuracy at 93.2% and cut the average time from 36 to 21 milliseconds. On posts, the calibrated confidence dropped, so almost nothing would be decided automatically. All of it ran on one workstation GPU, on AI the organisation owns.
Let's talk →
Can your model's confidence be trusted on your data? The answer is in your own examples. We are glad to measure it with you and help set up this pattern on AI your organisation owns. See On-Premise AI Platform and AI Consulting.
One 24 GB NVIDIA workstation GPU, CUDA 12.8; one request at a time. Times depend on the GPU: other cards will be faster or slower
Libraries
PyTorch 2.11, Hugging Face Transformers 5.17, PEFT 0.21
LLMs
Typhoon 2.5 4B (`typhoon-ai/typhoon2.5-qwen3-4b`), Qwen3 4B (`Qwen/Qwen3-4B`) and Qwen3 8B (`Qwen/Qwen3-8B`), in bfloat16. Each answer is read from the model's probabilities for the option letters, so it can only pick one of the listed options
Encoder
mmBERT-base (`jhu-clsp/mmBERT-base`), 32-bit, with a classification layer
Laya
`convaiinnovations/laya`, multilingual version, pinned revision, no extra training
Questions
Question and option descriptions written in English; the texts are Thai. In MASSIVE, spaces between Thai words were removed, as real Thai text has none
Fine-tuning
LLM: LoRA rank 16 on all attention and feed-forward layers, 3 epochs. Encoder: all weights, 6 epochs, learning rates chosen on a separate slice of the training data
Encoder in front
The encoder answers when its calibrated confidence is at least 90%, a threshold fixed in advance. Otherwise the request goes to the LLM: with all options (fine-tuned Typhoon), or with the encoder's top 3 of 8 topics or top 2 of 3 sentiments (untrained Qwen3 4B). Time per request is the encoder's time plus the LLM's time for the share it receives
English
Laya with no training, and Laya and mmBERT trained on about 150 English MASSIVE examples, tested on the English versions of the same 400 requests
Calibration
One temperature per model, fitted on a separate calibration set (200 MASSIVE and 150 Wisesight items per sample)
Confidence
Out of the box: each model's own probabilities (Laya ships temperature 1.0, so none; encoders and LLMs the softmax of their option scores). After one fit: one temperature fitted on the calibration set. A confident wrong answer: at least 0.9 on the chosen option, and wrong. Calibration error: expected calibration error, 10 bins. Cross check: for the models that answered both datasets as the same model, the temperature fitted on one dataset was applied to the other
Test sets
400 MASSIVE and 450 Wisesight Thai items per sample, balanced across options
Samples and statistics
3 random samples; accuracy differences checked with a paired bootstrap over test items (10,000 resamples, 95% intervals). It ignores variation from which training examples were drawn, so the intervals are optimistic. Calibration figures are means without intervals
Memory figure
Peak memory PyTorch allocated for one request, weights included. The GPU needs additional memory on top for the CUDA runtime and cache
Our measurements: 400 or 450 Thai test items per sample, 3 random samples, accuracy differences checked with a paired bootstrap, on one 24 GB NVIDIA workstation GPU
Thank you to Wisesight for making their dataset freely available to everyone.