Insight — AI Decisions in Thai

Jev Made "System One" AI Famous. In Thai, Confidence Is the Part to Check.

A System One model decides alone when its confidence is high. We tested this in Thai, on AI an organisation can own. Out of the box, the confidence was often wrong on social-media posts. One simple adjustment, made with a small set of labelled examples, made the confidence honest for every model we tested. Check the confidence on your own data before you let a model decide alone.

Infozense · CIO · Head of Data · ~18 minutes · Tested September 2026

Who this is for

This article is for CIOs, heads of data and IT leads in organisations that need AI to work in Thai, or in any language other than English. We use Thai as the test case. You may have seen the attention around Jev and "System One" models. You want the same benefit, and you want to know whether you can have it on AI your organisation owns or controls. This article reports what we measured when we tried the idea in Thai.

Jev makes the right point: most AI work is deciding, not writing

In September 2026 TypeSafe AI released a model called Jev. It is not designed for conversation. You send a question with a fixed list of answers, and Jev picks one answer and returns a probability for every option. TypeSafe says end-to-end response time is 70 to 500 milliseconds. TypeSafe calls this kind of model "System One", after the fast, intuitive thinking Daniel Kahneman describes in Thinking, Fast and Slow.

A figure titled “It picks from a fixed list and gives every option a probability”. A Thai request meaning “turn on the kitchen light”, and a fixed list of eight answers, go into a System One model. It returns a probability for every option, smart home 96.3% and alarm 0.9% at the top, and picks one answer: smart home, 96%. ภาพหัวข้อ “เลือกจากรายการที่กำหนด และให้ค่าความน่าจะเป็นกับทุกตัวเลือก” คำขอหนึ่งข้อ คือ “เปิดไฟในห้องครัว” พร้อมรายการคำตอบที่กำหนดไว้แปดข้อ ถูกส่งเข้าโมเดลแบบ System One โมเดลให้ค่าความน่าจะเป็นของทุกตัวเลือก โดยตัวเลือก บ้านอัจฉริยะ ได้ 96.3% และ นาฬิกาปลุก ได้ 0.9% แล้วเลือกคำตอบเดียวคือ บ้านอัจฉริยะ 96%

The idea behind it is right. Look at what organisations actually want from AI. Which department should handle this request? Is this complaint urgent? What kind of document is this? Each is a decision with a few possible answers. None of them needs a model that can write an essay.

What makes this work is the probability. TypeSafe says every Jev answer comes with calibrated probabilities and confidence scores. (Calibrated means that when the model says 90%, it is right about 90% of the time.) Reliable confidence lets the fast model answer on its own when the score is high enough, and send the rest to an LLM.

Jev is a managed service. TypeSafe runs it for you, and you pay for what you use: TypeSafe lists a price per million input tokens (a token is a small piece of text used to count usage). That makes it easy to start. Some organisations want more than that: they want to fine-tune the model on their own examples, or keep it on AI they own or control. For them there is another choice, and the rest of this article is about it.

What "System One" and "System Two" mean

In Thinking, Fast and Slow, the psychologist Daniel Kahneman describes two ways people think. He calls them System 1 and System 2.

System 1 is fast and automatic. You use it when you read a familiar word, recognise a face, or answer 2 + 2. It costs almost no effort, and it is right most of the time on familiar things.

System 2 is slow and deliberate. You use it when you multiply 17 by 24, or check a contract clause by clause. It takes effort and attention, so you use it only when you need to.

Most of your day runs on System 1. System 2 steps in when System 1 is unsure, or when something does not fit. That split is what makes people efficient: careful thinking is kept for the cases that need it.

AI can work the same way. An AI "System One" makes routine decisions fast, by picking from a fixed list of answers and saying how sure it is. An AI "System Two" is a slower, more capable model that takes the cases the fast part is unsure about. The rule that links them is confidence: when the fast part is sure, it answers; when it is not, it hands over. An organisation works the same way: a front desk handles routine requests and passes the unusual ones to a specialist.

A figure titled “The fast part decides most cases; the careful part takes the rest”. In people, System 1 is fast and automatic and System 2 is slow and deliberate. In AI you own, System One is an encoder of 1.2 GB that answers in about 11 ms, and System Two is a fine-tuned compact LLM of 7.6 GB that answers in about 36 ms and takes the cases the encoder is unsure about. ภาพหัวข้อ “ส่วนที่เร็วตัดสินใจเรื่องส่วนใหญ่ ส่วนที่รอบคอบรับเรื่องที่เหลือ” ฝั่งซ้ายคือการคิดของคน System 1 เร็วและอัตโนมัติ ส่วน System 2 ช้าแต่รอบคอบ ฝั่งขวาคือ AI ที่องค์กรเป็นเจ้าของ System One คือ encoder ขนาด 1.2 GB ตอบในราว 11 มิลลิวินาที ส่วน System Two คือ LLM กะทัดรัดที่ fine-tune แล้ว ขนาด 7.6 GB ตอบในราว 36 มิลลิวินาที และรับเฉพาะกรณีที่ค่า confidence ของ encoder ไม่ถึงเกณฑ์

TypeSafe named its model class after this idea. In this article we build both parts on AI an organisation owns: an encoder as System One, and a compact LLM as System Two.

Two kinds of language model: encoder and decoder. The fast part is an encoder. It reads the whole text in one pass: each word looks at the words before and after it, and keeps its position, so word order still counts. It gives each option a score, but it cannot write. The encoder we used is compact, 1.2 GB, and answers in about 11 milliseconds; that is a design choice, not something every encoder has. The careful part is a decoder, the kind of model behind chat AI and most LLMs. Each word looks only at the words before it, and the model writes by predicting the next word, one word at a time. In our test we asked it for a single word, the option letter, so its 36 milliseconds come from its larger size, 7.6 GB, and its longer input (the prompt carries the instructions and every option), not from writing long answers. The two are separate models. The encoder passes on only its answer and how sure it is, so either part can be replaced by a better one later. It is easy to assume that Jev is an LLM, but it is not a chat model: TypeSafe says it gives up text generation and returns structured answers. TypeSafe has not published how Jev is built. Laya, an open-source model that follows the same approach, is an encoder.

A figure titled “The encoder scores a fixed list; the decoder predicts the next word”. Two layer stacks side by side, each read from the bottom up. On the left, the encoder, System One: mmBERT-base, 308M parameters, 1.2 GB, about 11 ms. Its five layers are the input, the question, options and text all at once; an embedding where each token becomes 768 numbers, plus its position; 22 encoder layers whose attention looks BOTH ways, so each token sees the tokens before and after it; a compact scoring head giving one score per option; and the output, smart home 96%, alarm 1% and so on, where it picks and cannot write. On the right, the decoder, the kind of model behind LLMs, System Two: Typhoon 2.5 (Qwen3 4B), 4.0B parameters, 7.6 GB, about 36 ms for a one-word answer. The same five layers are the input, a prompt with the question, options and text; an embedding where each token becomes 2,560 numbers, plus its position; 36 decoder layers whose attention looks BACK only, so each token sees only the tokens before it; a vocabulary head giving a probability for each of its 151,936 tokens; and the output, the next token, in our test one run, reading only the option letters. A dashed loop runs from the decoder output back down to its input: to write, add the token and run again, token by token. Two separate models. The encoder passes on only its answer and how sure it is. Memory and time were measured one request at a time on one 24 GB workstation GPU. ภาพหัวข้อ “encoder ให้คะแนนตัวเลือกที่กำหนด ส่วน decoder ทำนายคำถัดไป” ภาพสองฝั่ง แต่ละฝั่งเป็นชั้นของโมเดลเรียงจากล่างขึ้นบน ฝั่งซ้ายคือ encoder (System One) ใช้ mmBERT-base มีพารามิเตอร์ 308M ใช้หน่วยความจำ 1.2 GB และตอบในราว 11 ms ชั้นล่างสุดคือข้อมูลนำเข้า ได้แก่ คำถาม ตัวเลือก และข้อความ ส่งเข้าพร้อมกัน ถัดขึ้นมาคือ embedding ที่แปลงแต่ละ token เป็นตัวเลข 768 ค่า พร้อมข้อมูลตำแหน่ง ถัดขึ้นไปคือ encoder 22 ชั้น ซึ่ง attention มองได้ทั้งสองทาง แต่ละ token จึงเห็นทั้ง token ก่อนหน้าและถัดไป แล้วส่งต่อให้ส่วนให้คะแนนแบบกะทัดรัด ซึ่งให้หนึ่งคะแนนต่อหนึ่งตัวเลือก ผลลัพธ์คือ บ้านอัจฉริยะ 96% นาฬิกาปลุก 1% และอื่น ๆ โดย encoder เลือกคำตอบได้ แต่เขียนข้อความไม่ได้ ฝั่งขวาคือ decoder ซึ่งเป็นโมเดลแบบที่อยู่เบื้องหลัง LLM (System Two) ใช้ Typhoon 2.5 (Qwen3 4B) มีพารามิเตอร์ 4.0B ใช้หน่วยความจำ 7.6 GB และตอบคำเดียวในราว 36 ms เรียงห้าชั้นแบบเดียวกัน คือข้อมูลนำเข้าเป็น prompt ที่มีคำถาม ตัวเลือก และข้อความ ถัดขึ้นมาคือ embedding ที่แปลงแต่ละ token เป็นตัวเลข 2,560 ค่า พร้อมข้อมูลตำแหน่ง ถัดขึ้นไปคือ decoder 36 ชั้น ซึ่ง attention มองได้เฉพาะย้อนหลัง แต่ละ token จึงเห็นเฉพาะ token ก่อนหน้า แล้วส่งต่อให้ส่วนทำนายคำ ซึ่งให้ความน่าจะเป็นของ token ทั้ง 151,936 ตัวในคลังคำ ผลลัพธ์คือ token ถัดไป ในการทดสอบนี้ ประมวลผลรอบเดียว อ่านเฉพาะตัวอักษรของตัวเลือก และมีเส้นประวนกลับจากผลลัพธ์ของ decoder ลงไปยังข้อมูลนำเข้า กำกับไว้ว่า ถ้าจะเขียนข้อความ ให้นำ token นั้นต่อท้าย แล้วประมวลผลใหม่ทีละ token ทั้งสองเป็นโมเดลแยกกันคนละตัว encoder ส่งต่อเฉพาะคำตอบและค่า confidence วัดหน่วยความจำและเวลาทีละคำขอ บน GPU แบบ workstation ขนาด 24 GB การ์ดเดียว

Why the confidence number is the part that matters

Everything in the System One idea hangs on one number: the confidence, the percentage that comes with each answer, such as "smart home 96%" in the first figure. If that confidence is at least 90%, the system uses the answer without checking it again. We use 90% in this article, but it is our choice, not a standard. A higher threshold lets fewer wrong answers through but sends more work to the LLM; the right level depends on what a wrong answer costs you. Whatever level you pick, check that the confidence is honest: for example, whether its 90% really means 90%.

If it does not, the routing goes wrong. A model that says 90% but is right only two times in three lets through one wrong answer in three, and nobody sees them. A model that says 60% when it is nearly always right hands over work it could have done itself.

Most published comparisons of System One models rank them by accuracy. We checked their confidence as well.

What we tested

Does a System One model's confidence hold up in Thai? And can the whole pattern run on AI your organisation owns or controls, without relying on someone else's service?

The short answers. Yes, it runs on AI you own: everything in this article ran on one GPU. But out of the box, no model's confidence held up on social-media posts. On routine requests every model passed the 90% check, though most were still somewhat overconfident. One cheap step handles both: adjust the confidence with a small set of labelled examples, and every model's confidence then matches how often it is right.

We tested six models:

  • An open-source System One model, with no training by us: Laya, from Convai Innovations, released in September 2026 under the Apache 2.0 licence.
  • Our own System One: an encoder (mmBERT) trained on 200 labelled examples.
  • Three LLMs with no training: Typhoon 2.5 4B, which is trained specifically for Thai, and the generic Qwen3 4B and Qwen3 8B. An LLM is a large language model that can read and write text; 4B means 4 billion parameters.
  • Typhoon 4B fine-tuned on the same 200 examples, as case one describes.

We tested them on a single 24 GB NVIDIA workstation GPU (the GPU is the processing card AI runs on). We used two public Thai datasets:

  • MASSIVE (Amazon, CC BY 4.0): short everyday requests to a voice assistant in 8 topics, such as "set an alarm for nine am" (alarm) or "can they provide takeaway" (takeaway). We tested on 400 Thai requests at a time.
  • Wisesight Sentiment (CC0): real Thai social-media posts in 3 sentiment groups. We tested on 450 posts at a time. This set is hard even for specialist models. WangchanBERTa is a Thai model widely used as a reference in Thai research. In its published paper it reached 76.19 micro-F1 after training on all of the roughly 21,600 posts. That work used 4 sentiment groups and a different metric from ours, so the two results cannot be compared directly.

Every result in this article is the mean of 3 random test samples. Accuracy differences were checked with a paired bootstrap over test items; the calibration figures are means without intervals. Their accuracy is in the table near the end. This article is about their confidence.

Out-of-the-box confidence failed on Thai posts

For each model we counted the answers it gave with at least 90% confidence, and how many of those were wrong. On social-media posts, out of the box, every model failed this check: of the answers given with at least 90% confidence, 25% to 47% were wrong, depending on the model. The chart in the next section shows a different number: confident wrong answers as a share of all posts. That share also depends on how often a model's confidence reached 90%. Laya's confidence reached 90% less often than the others', so its confident wrong answers came to 9.7% of all posts. Qwen3 4B was the worst: 45% of all posts got a confident wrong answer.

Training on the task did not prevent it. Our encoder was trained on 200 posts from the same dataset, and it was among the most overconfident: 27% of all posts got a confident wrong answer. Fine-tuning on few examples is known to make models overconfident.

On routine requests every model passed this check: of its answers given with at least 90% confidence, at least 90% were right. Most were still somewhat overconfident (calibration error 0.05 to 0.15), and between 3% and 9% of all requests got a confident wrong answer.

For the LLMs, we read the confidence from the probabilities of the option letters. That is what a developer would compute, not a number the vendor ships. This is one common way to read an LLM's confidence; others, such as using the whole vocabulary or changing the order of the options, can give different results.

One number made the confidence honest

The fix is cheap, and it worked on every model. It is called temperature scaling. You take 150 to 200 examples whose answers you know, and fit one number that softens the model's probabilities until its confidence matches how often it is right. The model and its answers stay the same. Only the confidence changes. The method comes from Guo et al. (2017), On Calibration of Modern Neural Networks, which also found that modern neural networks tend to be overconfident.

After one fit, for all six models, the share of all routine requests that got a confident wrong answer fell to at most 2.2%, and on social-media posts to almost zero. Look at how. On routine requests the models still reached 90% confidence on most requests (41% to 90% of them, depending on the model). On posts, almost no answer reached 90% any more (0% to 2.1% of posts), so almost nothing would be decided automatically. The calibration error, which measures the gap between confidence and accuracy, fell from between 0.23 and 0.47 to between 0.04 and 0.11 on social-media posts.

A chart titled “Confident but wrong: one fitted number made the confidence of all six honest”, with two panels and two bars per model: out of the box, and after one temperature fit on 150 to 200 labelled examples. The axis measures the share of ALL test items that were answered with at least 90% confidence and were wrong. On routine requests (MASSIVE), out of the box then after the fit: Laya 5.9% then 0.7%; our encoder trained on 200 examples 7.4% then 0.7%; Typhoon 4B 3.2% then 1.8%; Typhoon 4B fine-tuned 4.3% then 2.2%; Qwen3 4B 9.1% then 2.2%; Qwen3 8B 8.4% then 1.2%. On social-media posts (Wisesight): Laya 9.7% then 0.1%; our encoder 27.3% then 0.0%; Typhoon 4B 25.6% then 0.0%; Typhoon 4B fine-tuned 17.5% then 0.0%; Qwen3 4B 44.6% then 0.0%; Qwen3 8B 29.3% then 0.0%. The footnote states that the test texts were Thai, with the question and options in English; that the figures are the mean of 3 random samples; that the System One models are Laya and our encoder; that LLM confidence is read from the probabilities of the option letters; that our encoder and the fine-tuned Typhoon were trained on 200 examples of each dataset; that the others had no training; and that after the fit, the share of items still answered at 90% or more was 41 to 90% on routine requests and 0 to 2% on posts. กราฟหัวข้อ “ปรับ temperature เพียงค่าเดียว ค่า confidence ตรงกับความจริงทั้งหกโมเดล” แบ่งเป็นสองฝั่ง แต่ละโมเดลมีสองแท่ง คือ ใช้ทันทีโดยไม่ปรับ และหลังปรับ temperature หนึ่งค่า ด้วยตัวอย่างที่มีเฉลย 150 ถึง 200 ข้อ แกนของกราฟคือสัดส่วนของข้อทดสอบทั้งหมดที่ตอบผิด ทั้งที่มีค่า confidence ตั้งแต่ 90% ฝั่งคำขอในงานประจำ (MASSIVE) ก่อนปรับและหลังปรับ Laya 5.9% เป็น 0.7% encoder ของเราที่ฝึกด้วยตัวอย่าง 200 ข้อ 7.4% เป็น 0.7% Typhoon 4B 3.2% เป็น 1.8% Typhoon 4B ที่ fine-tune แล้ว 4.3% เป็น 2.2% Qwen3 4B 9.1% เป็น 2.2% และ Qwen3 8B 8.4% เป็น 1.2% ฝั่งโพสต์โซเชียลมีเดีย (Wisesight) Laya 9.7% เป็น 0.1% encoder ของเรา 27.3% เป็น 0.0% Typhoon 4B 25.6% เป็น 0.0% Typhoon 4B ที่ fine-tune แล้ว 17.5% เป็น 0.0% Qwen3 4B 44.6% เป็น 0.0% และ Qwen3 8B 29.3% เป็น 0.0% หมายเหตุใต้ภาพระบุว่า ข้อความทดสอบเป็นภาษาไทย คำถามและตัวเลือกเป็นภาษาอังกฤษ เป็นค่าเฉลี่ยจากการสุ่มทดสอบ 3 รอบ โมเดลแบบ System One คือ Laya และ encoder ของเรา ส่วนค่า confidence ของ LLM อ่านจากค่าความน่าจะเป็นของตัวอักษรตัวเลือก encoder และ Typhoon ที่ fine-tune แล้วของเรา ฝึกด้วยตัวอย่าง 200 ข้อจากแต่ละชุดข้อมูล โมเดลอื่นเราไม่ได้ฝึกเพิ่ม และหลังปรับ สัดส่วนข้อที่ค่า confidence ยังถึง 90% คือ งานประจำ 41 ถึง 90% และโพสต์ 0 ถึง 2%

Take our encoder on routine requests. The fit lowered the share of requests it was at least 90% sure of from 94% to 72%, and raised its accuracy on that share from 92% to 99%. The requests it was no longer sure of went to the LLM instead of slipping through.

After the fit, the models were almost never 90% sure of a social-media post (0% to 2.1% of posts), so at this threshold almost nothing would be decided automatically. That is the honest result for this task with 200 examples: none of the setups we tested is reliable enough to decide alone at 90%. Route those answers to a person, accept the accuracy of the best option (67.2% for the fine-tuned Typhoon 4B), or check accuracy at a lower threshold before ruling automation out.

Accuracy does not change with the fit. What changes is which answers you can let through unchecked.

Calibration holds only for the task and data it was fitted on

The fix has one condition: fit it on the same decision and the same kind of text the model will actually see. Laya, Typhoon and the two Qwen models answered both kinds of text as the very same model, so we could test this. We fitted the temperature on routine requests only, and then used that setting on social-media posts.

It did not carry over well. Calibrated on routine requests, Typhoon 4B was excellent there: only 1.8% of requests got a confident wrong answer. On social-media posts the same setting helped, but not enough: 15.7% of all posts still got a confident wrong answer (25.6% out of the box), and of its answers given with at least 90% confidence only 73% were right. Calibrated on the posts themselves, it gave no confident wrong answers at all.

It fails the other way too. Calibrated on social-media posts and used on routine requests, Typhoon was never 90% sure of anything, so it would have handed every routine request to the next step and thrown away the speed. The Qwen models behaved the same way, and Laya less strongly.

A chart titled “Fit on the wrong text: too sure on one, too unsure on the other”, with three bars per model: out of the box, temperature fitted on the OTHER kind of text, and temperature fitted on the SAME kind of text. The two panels measure different things. The left panel is routine requests, and its axis is the share answered with at least 90% confidence at all, which is what System One could take on alone: Laya 59%, 16%, 41%; Typhoon 4B 92%, 0%, 86%; Qwen3 4B 95%, 0%, 77%; Qwen3 8B 95%, 0%, 73%. So a temperature fitted on posts leaves three of the four models never confident on routine requests. The right panel is social-media posts, and its axis is the share of ALL posts answered with at least 90% confidence and wrong: Laya 9.7% out of the box, 1.6% with the other kind’s temperature, 0.1% with its own; Typhoon 4B 25.6%, 15.7%, 0.0%; Qwen3 4B 44.6%, 27.0%, 0.0%; Qwen3 8B 29.3%, 10.1%, 0.0%. The footnote states that the test texts were Thai, with the question and options in English; that each model answered both datasets as the very same model with no training; that a temperature fitted on routine requests was used on posts and one fitted on posts was used on routine requests; that it was fitted on 200 routine requests or 150 posts per sample; the mean of 3 random samples; and that LLM confidence is read from the option letters. กราฟหัวข้อ “ปรับผิดประเภท: ค่า confidence สูงเกินกับแบบหนึ่ง ต่ำเกินกับอีกแบบ” แต่ละโมเดลมีสามแท่ง คือ ใช้ทันทีโดยไม่ปรับ ปรับ temperature กับข้อความอีกประเภท และปรับ temperature กับข้อความประเภทเดียวกัน สองฝั่งของภาพวัดคนละอย่าง ฝั่งซ้ายคือคำขอในงานประจำ แกนคือสัดส่วนของคำขอที่ตอบด้วยค่า confidence ตั้งแต่ 90% ซึ่งเป็นส่วนที่ System One ตอบเองได้ ได้แก่ Laya 59% 16% และ 41% Typhoon 4B 92% 0% และ 86% Qwen3 4B 95% 0% และ 77% ส่วน Qwen3 8B 95% 0% และ 73% การปรับ temperature ด้วยโพสต์จึงทำให้ค่า confidence ของสามในสี่โมเดลไม่ถึงเกณฑ์เลยกับคำขอในงานประจำ ฝั่งขวาคือโพสต์โซเชียลมีเดีย แกนคือสัดส่วนของโพสต์ทั้งหมดที่ตอบผิด ทั้งที่มีค่า confidence ตั้งแต่ 90% ได้แก่ Laya 9.7% เมื่อไม่ปรับ 1.6% เมื่อปรับด้วยข้อความอีกประเภท และ 0.1% เมื่อปรับด้วยข้อความประเภทเดียวกัน Typhoon 4B 25.6% 15.7% และ 0.0% Qwen3 4B 44.6% 27.0% และ 0.0% ส่วน Qwen3 8B 29.3% 10.1% และ 0.0% หมายเหตุใต้ภาพระบุว่า ข้อความทดสอบเป็นภาษาไทย คำถามและตัวเลือกเป็นภาษาอังกฤษ แต่ละโมเดลใช้ตัวเดียวกันตอบทั้งสองชุดข้อมูล และไม่ได้ฝึกเพิ่ม โดยนำ temperature ที่ปรับกับคำขอในงานประจำไปใช้กับโพสต์ และทำกลับกันด้วย ปรับด้วยคำขอ 200 ข้อ หรือโพสต์ 150 ข้อ ต่อการสุ่มหนึ่งรอบ ค่าเฉลี่ยจากการสุ่มทดสอบ 3 รอบ และค่า confidence ของ LLM อ่านจากค่าความน่าจะเป็นของตัวอักษรตัวเลือก

So calibration belongs to the model and the decision together, not to the model alone. Our two datasets differ in task, number of options and difficulty as well as in text, so this test cannot say which of these matters most. The practical rule is the same: fit and check the temperature on the exact decision and data the model will see. Related research on shifted data points the same way: Ovadia et al. (2019), Can You Trust Your Model's Uncertainty?, found that calibration fitted on one kind of data falls short when the data shifts.

Two cases where a System One earns its place

Why pair a System One with an LLM at all? An LLM trained for Thai is accurate on its own: Typhoon 4B answered 91.5% of routine requests correctly with no training (our test; all results are in the table further down). A calibrated System One in front earns its place in two cases, one in each of the next two sections. On high-volume routine requests, it answers most of them itself, in about 11 milliseconds, and leaves the LLM free for the hard ones. And when you have to keep a generic LLM, it brings it level with a model trained for Thai.

Case one: with calibrated confidence, System One answers most routine requests itself

The System One idea does not have to live inside a single model. You can build it from two parts your organisation owns: a fast part that answers when it is sure, and a careful part for everything else.

System Two: a compact LLM, fine-tuned. We fine-tuned Typhoon 4B on 200 examples that already had their answers labelled. The technique was LoRA (a fine-tuning technique that trains only a small set of added parameters and leaves the original model untouched). On our 24 GB workstation GPU, training took 75 to 103 seconds and less than 10 GB of GPU memory. The encoder, trained on the same examples on the same card, took 12 to 20 seconds. Times will differ on other GPUs. The fine-tuned model answered 92.8% of routine requests correctly, and its 67.2% on social-media posts was the highest in our tests.

System One: an encoder in front. An encoder is a compact model that reads text and picks an answer from a list, without writing any text. Fine-tuned on the same 200 examples, the encoder used 1.2 GB of memory and answered in about 11 milliseconds. On its own it was right 89.6% of the time on routine requests and 60.0% on social-media posts. The fine-tuned 4B got 92.8% and 67.2%, with more than six times the memory and about three times the time. There is one rule. If the encoder is at least 90% confident, it answers itself. If not, the request goes on to System Two.

A figure titled “When the encoder is at least 90% sure, it answers; the rest go to the LLM”. Two requests reach System One, an encoder of 1.2 GB. For a Thai request meaning “turn on the kitchen light” it is 96% sure, smart home 96% against alarm 1%, and answers smart home itself in about 11 ms. For a Thai request meaning “call for take-out”, whose wording contains the Thai for “go home”, it is unsure: transport 48%, takeaway 36%, so its top pick is wrong. It hands over to System Two, a fine-tuned 4B LLM, which picks from the same list, with no sentence, and answers takeaway, 98%, in about 36 ms. ภาพหัวข้อ “ค่า confidence ของ encoder ถึง 90% ตอบเองได้ ไม่ถึง ส่งต่อให้ LLM” คำขอสองตัวอย่างเข้าสู่ System One ซึ่งเป็น encoder ขนาด 1.2 GB คำขอ “เปิดไฟในห้องครัว” ค่า confidence ของ encoder อยู่ที่ 96% (บ้านอัจฉริยะ 96% นาฬิกาปลุก 1%) encoder จึงตอบเองว่า บ้านอัจฉริยะ ในราว 11 มิลลิวินาที ส่วนคำขอ “โทรซื้อกลับบ้าน” ค่า confidence ของ encoder ไม่ถึงเกณฑ์ ได้ การเดินทาง 48% กับ สั่งอาหาร 36% ตัวเลือกอันดับแรกของ encoder จึงผิด encoder ส่งต่อให้ System Two ซึ่งเป็น LLM 4B ที่ fine-tune แล้ว ซึ่งเลือกจากรายการเดียวกัน ไม่ได้เขียนประโยค และตอบว่า สั่งอาหาร 98% ในราว 36 มิลลิวินาที

On routine requests the encoder answered 72% of requests itself. When it was that confident, it was right 98.3% to 99.6% of the time. Overall accuracy was 93.2%, with no significant difference from the fine-tuned LLM alone (92.8%). Estimated average response time, from the measured per-model times, fell from about 36 to about 21 milliseconds. Memory goes up instead. Both models stay loaded, so the pair needs 8.8 GB of GPU memory, against 7.6 GB for the fine-tuned LLM alone. What the encoder saves is time, not memory; only an encoder used on its own saves memory.

When text is varied, System One steps aside. On social-media posts the encoder reached 90% confidence on fewer than 1 post in 100. So it sent almost every post on to System Two. Accuracy stayed at 67.2%, but every post now went through both parts, so response time rose from 36.5 milliseconds for the fine-tuned LLM alone to 47.2 milliseconds.

A chart titled “Same accuracy; faster only on routine requests”, with two panels. The bars are response time, because the accuracy is the same. On routine requests, the fine-tuned 4B LLM alone takes 36 ms at 92.8% accuracy, and the encoder in front of the fine-tuned 4B takes 21 ms at 93.2%, with the encoder answering 72% of requests itself. On social-media posts, the 4B alone takes 36 ms at 67.2%, and the encoder in front takes 47 ms at the same 67.2%, with the encoder answering under 1% of requests itself: putting it in front costs time there instead of saving it. กราฟหัวข้อ “แม่นเท่าเดิม แต่เร็วขึ้นเฉพาะคำขอในงานประจำ” แบ่งเป็นสองฝั่ง ความยาวของแท่งคือเวลาตอบ เพราะความแม่นยำเท่ากัน ฝั่งคำขอในงานประจำ LLM 4B ที่ fine-tune แล้วตัวเดียว ใช้เวลา 36 ms ที่ความแม่นยำ 92.8% ส่วน encoder วางไว้ด้านหน้า LLM 4B ที่ fine-tune แล้ว ใช้เวลา 21 ms ที่ความแม่นยำ 93.2% โดย encoder ตอบเอง 72% ของคำขอ ฝั่งโพสต์โซเชียลมีเดีย LLM 4B ตัวเดียวใช้เวลา 36 ms ที่ความแม่นยำ 67.2% ส่วน encoder วางไว้ด้านหน้า ใช้เวลา 47 ms ที่ความแม่นยำ 67.2% เท่ากัน โดย encoder ตอบเองไม่ถึง 1% ของคำขอ การวาง encoder ไว้ด้านหน้าในงานแบบนี้จึงเพิ่มเวลา ไม่ได้ลดเวลา

That is a useful signal. The encoder's confidence is calibrated, so if System One is almost never confident on your data, the task is too hard for it with the examples you have. Let the fine-tuned LLM decide on its own.

Case two: when you must keep a generic LLM, a Thai encoder in front brings it level

Why not simply use Typhoon? If you can choose, do: a model trained for Thai on its own is the simpler answer. But many organisations already run a generic LLM such as Qwen for other work, and have tested, approved and secured that one model. Adding a model trained for Thai would mean running a second 4B LLM: about 7.6 GB more GPU memory, and one more model to test, approve and update. An encoder adds about 1.2 GB. Can it close the Thai gap instead?

We tested this with Qwen3 at 4B, the generic model that Typhoon 2.5 is built on, with no extra training. On its own it answered 87.3% of routine requests and 52.1% of social-media posts correctly. That is 4.2 and 10.6 points below Typhoon, and both gaps are outside statistical noise. Same size, same design: what differs is Typhoon's additional training for Thai.

Then we put the encoder from case one in front of it, trained on the same 200 Thai examples. Qwen itself was not trained at all. The encoder does two jobs. When it is at least 90% sure, it answers itself. When it is not, it gives Qwen a shortlist of its most likely options, 3 of the 8 topics or 2 of the 3 sentiments, and Qwen picks from that shortlist.

With the encoder in front, Qwen3 4B reached 91.9% on routine requests and 61.7% on social-media posts. That is level with Typhoon alone, at 91.5% and 62.7%. Neither difference is statistically clear: the 95% confidence intervals are −1.1 to +1.9 points and −3.4 to +1.5 points. Note that the pair used 200 labelled examples and untrained Typhoon none; Typhoon fine-tuned on the same examples reached 67.2% on posts. So the encoder does not make a generic model better than one trained for Thai. It makes it as good, without changing the generic model.

A chart titled “If you must keep a generic LLM: a Thai encoder in front brings it level with the Thai-trained LLM”, with two panels. On routine requests (MASSIVE, 8 topics): a generic 4B LLM alone (Qwen3 4B) 87.3% at 35 ms and 7.5 GB; a Thai encoder in front of the same Qwen3 4B 91.9% at 21 ms and 8.7 GB; a Thai 4B LLM alone (Typhoon 2.5) 91.5% at 36 ms and 7.6 GB. On social-media posts (Wisesight, 3 sentiments): 52.1% at 34 ms, 61.7% at 45 ms, and 62.7% at 36 ms. The footnote states that no LLM was fine-tuned; that the encoder was trained on 200 Thai examples and answers when it is at least 90% sure, otherwise Qwen picks from the encoder’s top options; and that between the encoder in front and Typhoon there is no statistically clear difference on either dataset. Mean of 3 random samples on one 24 GB workstation GPU; times differ on other GPUs. กราฟหัวข้อ “ถ้าต้องใช้ LLM ทั่วไปต่อไป วาง encoder ภาษาไทยไว้ด้านหน้า ก็แม่นเท่า LLM ที่ฝึกมาเพื่อภาษาไทย” แบ่งเป็นสองฝั่ง ฝั่งคำขอในงานประจำ (MASSIVE 8 หัวข้อ) LLM ทั่วไป 4B ตัวเดียว (Qwen3 4B) ได้ 87.3% ใช้เวลา 35 ms และหน่วยความจำ 7.5 GB encoder ภาษาไทย วางไว้ด้านหน้า Qwen3 4B ตัวเดิม ได้ 91.9% ใช้เวลา 21 ms และ 8.7 GB ส่วน LLM ภาษาไทย 4B ตัวเดียว (Typhoon 2.5) ได้ 91.5% ใช้เวลา 36 ms และ 7.6 GB ฝั่งโพสต์โซเชียลมีเดีย (Wisesight 3 กลุ่มความรู้สึก) ได้ 52.1% ใช้เวลา 34 ms 61.7% ใช้เวลา 45 ms และ 62.7% ใช้เวลา 36 ms ตามลำดับ หมายเหตุใต้ภาพระบุว่า LLM ทุกตัวในแผนภูมินี้ไม่ได้ fine-tune ส่วน encoder ฝึกด้วยตัวอย่างภาษาไทย 200 ข้อ ถ้าค่า confidence ของ encoder อยู่ที่ 90% ขึ้นไป encoder จะตอบเอง ถ้าไม่ถึง Qwen จะเลือกจากตัวเลือกอันดับต้นที่ encoder คัดมาให้ และเมื่อใช้ encoder ความต่างจาก Typhoon ไม่มีนัยสำคัญทางสถิติในทั้งสองชุดข้อมูล ค่าเฉลี่ยจากการสุ่มทดสอบ 3 รอบ บน GPU แบบ workstation ขนาด 24 GB การ์ดเดียว และเวลาจะต่างไปตาม GPU ที่ใช้

On routine requests the pair was also faster, because the encoder answered most requests itself: about 21 milliseconds on average, against about 36 for Typhoon alone. On social-media posts the shortlist did the work. The 90% rule alone added almost nothing (52.2%), because the encoder was rarely that sure. Every post went through both parts, so the pair was slower there than Typhoon alone: 45 milliseconds against 36. The pair needs 8.7 GB of GPU memory, against 7.6 GB for Typhoon alone.

The practical point is that the encoder does not change Qwen at all. The same model keeps serving the rest of your work. When a new version of Qwen comes out, you can swap it in without retraining the encoder. You only measure again.

Every result in one table

StageOptionRoutine requests (MASSIVE)Social-media posts (Wisesight)Time per requestGPU memory
No extra trainingLaya (open-source System One model)68.9%51.1%about 13 ms / 12 ms1.2 GB
Generic 8B LLM (Qwen3 8B)87.5%62.3%57.6 ms / 38.9 ms15.3 GB
Generic 4B LLM (Qwen3 4B)87.3%52.1%34.7 ms / 34.5 ms7.5 GB / 7.6 GB
Thai 4B LLM (Typhoon 2.5)91.5%62.7%35.7 ms / 35.9 ms7.6 GB
Trained on 200 examplesEncoder (mmBERT)89.6%60.0%10.7 ms / 10.9 ms1.2 GB
Thai 4B LLM, fine-tuned with LoRA92.8%67.2%36.1 ms / 36.5 ms7.6 GB
Both togetherEncoder in front of the fine-tuned 4B LLM93.2%67.2%21.0 ms / 47.2 ms8.8 GB
Encoder in front of the generic 4B, untrained (shortlist)91.9%61.7%20.6 ms / 45.2 ms8.7 GB

Mean of 3 random test samples. Time measured one request at a time on one 24 GB workstation GPU; times will differ on other GPUs. Where a cell has two values, the first is routine requests and the second is social-media posts. GPU memory is the peak memory PyTorch allocated for one request, model weights included; a GPU needs additional memory on top of this for the CUDA runtime and PyTorch's cache, so size hardware with headroom. The encoder ran in 32-bit precision, the LLMs in 16-bit.

Before you let any System One decide alone

Whichever model you choose, open or managed, check it on your own work first:

  1. Collect labelled examples of the decision from your real work, in three separate sets: about 200 to train on if you fine-tune, about 200 to fit the confidence, and a few hundred more to check the result. That is roughly 600 or more in total.
  2. Measure accuracy on the check set.
  3. Check the confidence on the check set. Of the answers given with at least 90% confidence, count how many are wrong.
  4. If too many are wrong, fit one temperature on the second set, never on the training examples. It changes no answers.
  5. Set the threshold on the calibrated number, and check it on the check set. Above it, System One answers. Below it, an LLM or a person decides. Look at two numbers: how many answers stay above the threshold, and how many of those are wrong.
  6. Check again when your data changes, or when you change the model.

Which AI setup fits your work: a decision map

Put together, our results give a short decision map. Every number in it comes from our tests. The map shows the main paths; item 2 below is a cheaper variant of item 5.

A decision map titled “Three questions pick the setup: language, LLM choice, volume”, starting from a decision task: pick one answer from a fixed list. First question, is the text in English? If yes, an encoder alone may be enough: Laya with no training 85.3%, encoders trained on about 150 examples 93 to 94%, 1.2 GB of GPU memory, with the note that this measured accuracy only, on routine requests, and that we did not test confidence or LLMs in English. If the text is Thai or another language, the next question is whether you can choose which LLM to run. If a generic LLM is fixed, put a Thai encoder in front of that LLM: Qwen3 4B goes from 87.3% to 91.9% and from 52.1% to 61.7%, level with the Thai-trained Typhoon and with Qwen itself unchanged, although on varied text it is slower than Typhoon, 45 against 36 ms. If you can choose, use a compact LLM trained for your language: Typhoon 2.5 4B with no training 91.5% and 62.7%, fine-tuned on 200 examples 92.8% and 67.2%. Then, are there many routine requests? If yes, add the encoder in front as System One, trained on the same 200 examples: it answers 72% of requests itself, average time falls from 36 to 21 ms, and memory rises from 7.6 to 8.8 GB. If the text is varied instead, keep the fine-tuned LLM alone, because the encoder is rarely sure and so only adds time, from 36.5 to 47.2 ms. The two numbers throughout are routine requests (MASSIVE) and social-media posts (Wisesight). Thai results are the mean of 3 random samples on one 24 GB workstation GPU; times differ on other GPUs; every part runs on AI your organisation owns or controls; and whichever branch you take, measure on your own examples. ผังการตัดสินใจ หัวข้อ “สามคำถามเลือกรูปแบบ: ภาษา การเลือก LLM และปริมาณงาน” เริ่มจากงานตัดสินใจ คือเลือกหนึ่งคำตอบจากรายการที่กำหนด คำถามแรกคือ ข้อความเป็นภาษาอังกฤษหรือไม่ ถ้าใช่ encoder ตัวเดียวอาจเพียงพอ โดย Laya ที่ไม่ฝึกเพิ่มได้ 85.3% ส่วน encoder ที่ฝึกด้วยตัวอย่างที่มีเฉลยราว 150 ข้อ ได้ 93 ถึง 94% ใช้หน่วยความจำ GPU 1.2 GB พร้อมหมายเหตุว่าวัดเฉพาะความแม่นยำของคำขอในงานประจำ และเราไม่ได้ทดสอบค่า confidence และ LLM ในภาษาอังกฤษ ถ้าไม่ใช่ คือเป็นภาษาไทยหรือภาษาอื่น คำถามถัดไปคือ เลือก LLM เองได้หรือไม่ ถ้าเลือกไม่ได้ เพราะต้องใช้ LLM ทั่วไปที่กำหนดไว้ ให้วาง encoder ภาษาไทยไว้ด้านหน้า LLM ตัวนั้น Qwen3 4B จะขึ้นจาก 87.3% เป็น 91.9% และจาก 52.1% เป็น 61.7% ซึ่งแม่นเท่า Typhoon ที่ฝึกมาเพื่อภาษาไทย โดยไม่เปลี่ยนแปลง Qwen แต่กับข้อความหลากหลายจะช้ากว่า Typhoon คือ 45 เทียบกับ 36 ms ถ้าเลือกได้ ให้ใช้ LLM กะทัดรัดที่ฝึกมาเพื่อภาษาของคุณ โดย Typhoon 2.5 4B ที่ไม่ฝึกเพิ่มได้ 91.5% และ 62.7% ส่วนที่ fine-tune ด้วยตัวอย่างที่มีเฉลย 200 ข้อ ได้ 92.8% และ 67.2% จากนั้นถามว่ามีคำขอในงานประจำจำนวนมากหรือไม่ ถ้าใช่ ให้วาง encoder ไว้ด้านหน้าให้เป็น System One โดยฝึกด้วยตัวอย่างที่มีเฉลย 200 ข้อชุดเดียวกัน encoder จะตอบเอง 72% ของคำขอ เวลาเฉลี่ยลดจาก 36 เหลือ 21 ms และหน่วยความจำเพิ่มจาก 7.6 เป็น 8.8 GB ถ้าเป็นข้อความหลากหลาย ให้ใช้ LLM ที่ fine-tune แล้วตัวเดียว เพราะค่า confidence ของ encoder แทบไม่ถึงเกณฑ์ จึงมีแต่เพิ่มเวลา จาก 36.5 เป็น 47.2 ms ตัวเลขสองค่าในผังคือ คำขอในงานประจำ (MASSIVE) และโพสต์โซเชียลมีเดีย (Wisesight) ผลในภาษาไทยเป็นค่าเฉลี่ยจากการสุ่มทดสอบ 3 รอบ บน GPU แบบ workstation ขนาด 24 GB การ์ดเดียว เวลาจะต่างไปตาม GPU ที่ใช้ ทุกส่วนทำงานบน AI ที่องค์กรเป็นเจ้าของหรือควบคุมเอง และไม่ว่าจะเลือกทางใด คุณควรวัดผลกับตัวอย่างงานของคุณเอง
  1. If your text is in English, an encoder on its own may be enough. With no training, Laya answered 85.3% of the English routine requests correctly. Encoders trained on about 150 English examples reached 93.2% to 93.9%, with 1.2 GB of GPU memory. We did not test LLMs in English, so this branch compares encoders only.
  2. If about 90% accuracy is enough, or your hardware is tight, an encoder trained on your examples can decide alone. On routine requests ours was right 89.6% of the time, in about 11 milliseconds with 1.2 GB of GPU memory. Or let it answer only when its confidence is 90% or more: it then answered 72% of routine requests, right about 99% of the time, and flagged the other 28% for a person to check.
  3. If your text is in Thai or another language, and you can choose the LLM, use a compact LLM trained for that language. Then fine-tune it on about 200 examples of your organisation's decision. On one workstation GPU this took under two minutes.
  4. If you must keep a generic LLM, put an encoder trained on your language in front of it. In our test this brought Qwen3 4B level with Typhoon, without changing Qwen.
  5. If requests are high in volume and routine, put an encoder in front of the fine-tuned LLM. That encoder is your System One. It answered 72% of routine requests itself and cut the average time from 36 to 21 milliseconds.
  6. If the text is varied, let the fine-tuned LLM decide alone. Calibrated confidence tells you which case you have: if System One is almost never sure on your data, it only adds time.
  7. Keep every part on AI your organisation owns or controls. Every branch runs on a single 24 GB GPU.

Limits of this test

  • We did not test Jev, and we do not claim to be more accurate than Jev.
  • Our main results are in Thai only. The pattern should carry over to other languages beyond English, but the numbers will not. Measure in your own language.
  • Laya's developers do not publish its training data, so it may have seen the MASSIVE requests during training. That would flatter its English score most.
  • In English we tested encoders only, on routine requests. We measured their accuracy, not their confidence. We did not test LLMs in English.
  • We did not fine-tune the generic Qwen3 4B, so we cannot say how a fine-tuned generic model would compare with an encoder in front of it. We also did not test Qwen's other abilities: "does not change Qwen" means its weights are not touched.
  • We tested two public datasets. The texts are short, single messages, not conversations.
  • Our 200 examples came from those datasets' training sets, not from your data. Your work will give different results, and the only way to know is to measure on your own examples.
  • We timed one request at a time in plain PyTorch (PyTorch is the standard library for running AI models) on one workstation GPU. Times depend on the GPU: other cards will be faster or slower, and on a tuned production server every option would likely be faster.
  • Every result is the mean of 3 random samples. Trained models vary from sample to sample, depending on which examples they get, and our significance tests do not fully capture that. So we report close comparisons as no clear difference, never as a win.
  • For the LLMs, out-of-the-box confidence is our reading of the option-letter probabilities. Their vendors do not ship it as a calibrated number.
  • We did not test Jev's confidence. TypeSafe says its probabilities are calibrated; check them on your own data, as with any model.
  • Our test sets were balanced across options. Real work is not, and calibration depends on the mix of cases. Fit and check on a sample of your real work.
  • What we present is a pattern and a measurement, not a ready-made service.
  • These results hold for the six models we tested. In a later test, a frozen embedding model with a simple classifier on top was not overconfident out of the box. Check whichever model you use.

The bottom line

Jev is right that most AI work is deciding, not writing, and that a fast model can decide alone when its confidence can be trusted. In Thai, out of the box, the confidence of all six models we tested failed on social-media posts: their confident answers were wrong a quarter to almost half of the time, even for a model trained on the task. One number, fitted on 150 to 200 labelled examples of the same kind of text, made the confidence of all six honest. With honest confidence, a System One in front of an LLM answered 72% of routine requests itself, kept accuracy at 93.2% and cut the average time from 36 to 21 milliseconds. On posts, the calibrated confidence dropped, so almost nothing would be decided automatically. All of it ran on one workstation GPU, on AI the organisation owns.

Let's talk →

Can your model's confidence be trusted on your data? The answer is in your own examples. We are glad to measure it with you and help set up this pattern on AI your organisation owns. See On-Premise AI Platform and AI Consulting.

Start a Conversation

contact@infozense.com  |  +66-82-242-4008  |  Bangkok, Thailand

How we measured

GPUOne 24 GB NVIDIA workstation GPU, CUDA 12.8; one request at a time. Times depend on the GPU: other cards will be faster or slower
LibrariesPyTorch 2.11, Hugging Face Transformers 5.17, PEFT 0.21
LLMsTyphoon 2.5 4B (`typhoon-ai/typhoon2.5-qwen3-4b`), Qwen3 4B (`Qwen/Qwen3-4B`) and Qwen3 8B (`Qwen/Qwen3-8B`), in bfloat16. Each answer is read from the model's probabilities for the option letters, so it can only pick one of the listed options
EncodermmBERT-base (`jhu-clsp/mmBERT-base`), 32-bit, with a classification layer
Laya`convaiinnovations/laya`, multilingual version, pinned revision, no extra training
QuestionsQuestion and option descriptions written in English; the texts are Thai. In MASSIVE, spaces between Thai words were removed, as real Thai text has none
Fine-tuningLLM: LoRA rank 16 on all attention and feed-forward layers, 3 epochs. Encoder: all weights, 6 epochs, learning rates chosen on a separate slice of the training data
Encoder in frontThe encoder answers when its calibrated confidence is at least 90%, a threshold fixed in advance. Otherwise the request goes to the LLM: with all options (fine-tuned Typhoon), or with the encoder's top 3 of 8 topics or top 2 of 3 sentiments (untrained Qwen3 4B). Time per request is the encoder's time plus the LLM's time for the share it receives
EnglishLaya with no training, and Laya and mmBERT trained on about 150 English MASSIVE examples, tested on the English versions of the same 400 requests
CalibrationOne temperature per model, fitted on a separate calibration set (200 MASSIVE and 150 Wisesight items per sample)
ConfidenceOut of the box: each model's own probabilities (Laya ships temperature 1.0, so none; encoders and LLMs the softmax of their option scores). After one fit: one temperature fitted on the calibration set. A confident wrong answer: at least 0.9 on the chosen option, and wrong. Calibration error: expected calibration error, 10 bins. Cross check: for the models that answered both datasets as the same model, the temperature fitted on one dataset was applied to the other
Test sets400 MASSIVE and 450 Wisesight Thai items per sample, balanced across options
Samples and statistics3 random samples; accuracy differences checked with a paired bootstrap over test items (10,000 resamples, 95% intervals). It ignores variation from which training examples were drawn, so the intervals are optimistic. Calibration figures are means without intervals
Memory figurePeak memory PyTorch allocated for one request, weights included. The GPU needs additional memory on top for the CUDA runtime and cache

Sources

Thank you to Wisesight for making their dataset freely available to everyone.