OpenThai-SystemOne

OpenThai-SystemOne: an open Thai + English System One decision model, 0.8B, Apache-2.0

An open Thai + English "System One" decision model. It does not generate text. Given a state (any text or JSON) and typed questions, it returns calibrated probabilities over the options in one forward pass:

Question You give You get
choice instructions + up to 255 options (name โ†’ description or null) choice, probabilities, confidence
score instructions + 2โ€“10 ordered level descriptions score (probability-weighted, fractional), probabilities, confidence
noul a yes/no question noul = p(yes)

The request/response contract mirrors TypeSafe's POST /v1/systemone so code written for the TypeSafe SDK can be pointed at this model unchanged. Typical uses: ticket routing, moderation, intent detection, RAG relevance judging, LLM-output verification, and computer-use / browser-agent action selection (which element to click, which tool to call).

How it works

  • Backbone: text tower of Qwen/Qwen3.5-0.8B-Base (24 layers, hybrid Gated-DeltaNet / attention, 262k context), vision encoder removed, then continued-pretrained on ~5B tokens of Thai (web, Wikipedia, parallel Thaiโ†”English, and machine-state text such as accessibility trees and JSON).
  • The 248k-token LM head is replaced by a 256-way slot head. Options are introduced by control tokens <|ts_opt_0|> โ€ฆ <|ts_opt_254|>; the hidden state at each <|ts_answer|> token is projected to 256 logits, slots beyond the number of options are masked, and a softmax gives the distribution. Slot 255 is abstain (none of the options fit).
  • Trained on ~2โ€“3M decision examples converted from public Thai/English classification, NLI, QA, rating, agent and tool-selection datasets plus synthetic Thai/English decision tasks, with option-order shuffling and abstain examples; then a short calibration stage (Brier loss + per-type temperature) so that higher confidence โ‡’ higher accuracy.

Usage

pip install openthai-systemone            # or: pip install "git+https://github.com/iapp-technology/openthai-systemone"
from openthai_systemone import SystemOneClient, Choice, Score, Noul

client = SystemOneClient("iapp/OpenThai-SystemOne")
resp = client.system_one(
    state={"ticket": "เธฅเธนเธเธ„เน‰เธฒเนเธˆเน‰เธ‡เธงเนˆเธฒเน‚เธ”เธ™เธซเธฑเธเน€เธ‡เธดเธ™เธ‹เน‰เธณเธชเธญเธ‡เธ„เธฃเธฑเน‰เธ‡ เธ‚เธญเน€เธ‡เธดเธ™เธ„เธทเธ™เธ”เนˆเธงเธ™ เน‚เธ—เธฃเธกเธฒเธชเธฒเธกเธฃเธญเธšเนเธฅเน‰เธง"},
    questions={
        "department": Choice(instructions="เธ—เธตเธกเนƒเธ”เธ„เธงเธฃเธฃเธฑเธšเธœเธดเธ”เธŠเธญเธš", criteria={"billing": "เธเธฒเธฃเน€เธ‡เธดเธ™/เธ„เนˆเธฒเธšเธฃเธดเธเธฒเธฃ", "technical": "เธฃเธฐเธšเธšเนƒเธŠเน‰เธ‡เธฒเธ™เน„เธกเนˆเน„เธ”เน‰", "sales": None}),
        "frustration": Score(instructions="เธฅเธนเธเธ„เน‰เธฒเธซเธ‡เธธเธ”เธซเธ‡เธดเธ”เนเธ„เนˆเน„เธซเธ™", criteria=["เนƒเธˆเน€เธขเน‡เธ™", "เธซเธ‡เธธเธ”เธซเธ‡เธดเธ”เนเธ•เนˆเธชเธธเธ เธฒเธž", "เน‚เธเธฃเธ˜เธกเธฒเธ"]),
        "refund_requested": Noul(instructions="เธฅเธนเธเธ„เน‰เธฒเธ‚เธญเน€เธ‡เธดเธ™เธ„เธทเธ™เธญเธขเนˆเธฒเธ‡เธŠเธฑเธ”เน€เธˆเธ™เธซเธฃเธทเธญเน„เธกเนˆ"),
    },
)
print(resp.answers["department"].choice, resp.answers["department"].probabilities)
print(resp.answers["frustration"].score, resp.answers["refund_requested"].noul)

HTTP server with the TypeSafe-compatible contract:

OPENTHAI_SYSTEMONE_MODEL=iapp/OpenThai-SystemOne uvicorn openthai_systemone.server:app --port 8000
curl -X POST localhost:8000/v1/systemone -H 'content-type: application/json' -d '{"state": "...", "questions": {...}}'

Or with plain transformers (remote code shipped in this repo):

from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("iapp/OpenThai-SystemOne", trust_remote_code=True)

Ollama (โ‰ฅ 0.35)

Ollama 0.35 serves System One models locally at POST /v1/systemone, with the same Jev-style contract. OpenThai-SystemOne has an Ollama build: iapp/openthai-systemone on ollama.com (tags 0.8b = 0.8b-q8_0, 0.8b-q4_K_M, 0.8b-bf16), also as GGUF on Hugging Face: iapp/OpenThai-SystemOne-Ollama. No API key; nothing leaves your machine.

ollama pull iapp/openthai-systemone          # or: ollama pull hf.co/iapp/OpenThai-SystemOne-Ollama:Q8_0
curl http://localhost:11434/v1/systemone -d '{
  "model": "iapp/openthai-systemone",
  "state": {"ticket": "เธฅเธนเธเธ„เน‰เธฒเนเธˆเน‰เธ‡เธงเนˆเธฒเน‚เธ”เธ™เธซเธฑเธเน€เธ‡เธดเธ™เธ‹เน‰เธณเธชเธญเธ‡เธ„เธฃเธฑเน‰เธ‡ เธ‚เธญเน€เธ‡เธดเธ™เธ„เธทเธ™เธ”เนˆเธงเธ™ เน‚เธ—เธฃเธกเธฒเธชเธฒเธกเธฃเธญเธšเนเธฅเน‰เธง"},
  "questions": {
    "department": {"type": "choice", "instructions": "เธ—เธตเธกเนƒเธ”เธ„เธงเธฃเธฃเธฑเธšเธœเธดเธ”เธŠเธญเธš",
                   "criteria": {"billing": "เธเธฒเธฃเน€เธ‡เธดเธ™/เธ„เนˆเธฒเธšเธฃเธดเธเธฒเธฃ", "technical": "เธฃเธฐเธšเธšเนƒเธŠเน‰เธ‡เธฒเธ™เน„เธกเนˆเน„เธ”เน‰", "sales": null}},
    "frustration": {"type": "score", "instructions": "เธฅเธนเธเธ„เน‰เธฒเธซเธ‡เธธเธ”เธซเธ‡เธดเธ”เนเธ„เนˆเน„เธซเธ™", "criteria": ["เนƒเธˆเน€เธขเน‡เธ™", "เธซเธ‡เธธเธ”เธซเธ‡เธดเธ”เนเธ•เนˆเธชเธธเธ เธฒเธž", "เน‚เธเธฃเธ˜เธกเธฒเธ"]},
    "refund_requested": {"type": "noul", "instructions": "เธฅเธนเธเธ„เน‰เธฒเธ‚เธญเน€เธ‡เธดเธ™เธ„เธทเธ™เธญเธขเนˆเธฒเธ‡เธŠเธฑเธ”เน€เธˆเธ™เธซเธฃเธทเธญเน„เธกเนˆ"}
  }
}'

Answer (Q8_0, probabilities rounded):

{"model": "iapp/openthai-systemone",
 "answers": {
   "department": {"type": "choice", "choice": "billing", "probabilities": {"billing": 0.9515, "technical": 0.0334, "sales": 0.0151}, "confidence": 0.7959},
   "frustration": {"type": "score", "score": 1.8771, "legend": {"0": "เนƒเธˆเน€เธขเน‡เธ™", "1": "เธซเธ‡เธธเธ”เธซเธ‡เธดเธ”เนเธ•เนˆเธชเธธเธ เธฒเธž", "2": "เน‚เธเธฃเธ˜เธกเธฒเธ"},
                   "probabilities": {"0": 0.0216, "1": 0.0797, "2": 0.8987}, "confidence": 0.6537},
   "refund_requested": {"type": "noul", "noul": 0.9805}},
 "usage": {"input_tokens": 741, "output_tokens": 4}}

Ollama does not run this repo's 256-slot head: it asks each question as a chat prompt and reads the probabilities of the answer letters Aโ€“Z. The Ollama build is therefore v0.3 fine-tuned for Ollama's prompt (same training data; accuracy through Ollama equals this repo's API, see Evaluation). Through Ollama a question takes 2โ€“26 options (this repo's API: 255), and there is no order-invariant mode and no abstain answer.

Demos

Thai showcase (real v0.3 outputs via the API)

OpenThai-SystemOne Thai showcase: 7 everyday Thai tasks with the model's real probabilities

Seven everyday Thai tasks, each answered in one forward pass; probabilities are the model's actual output. Example 1 keeps a real miss (politeness) on purpose.

Full inputs, questions and outputs of the 7 examples

1. Support ticket triage โ€” state: "เนเธญเธ›เน‚เธญเธ™เน€เธ‡เธดเธ™เน„เธกเนˆเน„เธ”เน‰เธ•เธฑเน‰เธ‡เนเธ•เนˆเน€เธกเธทเนˆเธญเธ„เธทเธ™ เธ‚เธถเน‰เธ™เธงเนˆเธฒ error 502 เธ•เธฅเธญเธ” เธฅเธญเธ‡เธฅเธ‡เนƒเธซเธกเนˆเนเธฅเน‰เธงเธเน‡เธขเธฑเธ‡เน„เธกเนˆเธซเธฒเธข เธฃเธšเธเธงเธ™เธŠเนˆเธงเธขเธ”เนˆเธงเธ™เธ™เธฐเธ„เธฃเธฑเธš เธ•เน‰เธญเธ‡เน‚เธญเธ™เธ„เนˆเธฒเน€เธ—เธญเธกเธฅเธนเธเธžเธฃเธธเนˆเธ‡เธ™เธตเน‰"

question type answer
เธ—เธตเธกเนƒเธ”เธ„เธงเธฃเธฃเธฑเธšเธœเธดเธ”เธŠเธญเธš (billing / technical / sales / account) choice technical 72%, billing 27%
เธ„เธงเธฒเธกเน€เธฃเนˆเธ‡เธ”เนˆเธงเธ™ (เน„เธกเนˆเน€เธฃเนˆเธ‡เธ”เนˆเธงเธ™ โ†’ เธงเธดเธเธคเธ•) score 2.13 = เน€เธฃเนˆเธ‡เธ”เนˆเธงเธ™ (85%), เธงเธดเธเธคเธ• 14%
เธฅเธนเธเธ„เน‰เธฒเนƒเธŠเน‰เธ–เน‰เธญเธขเธ„เธณเธชเธธเธ เธฒเธžเธซเธฃเธทเธญเน„เธกเนˆ noul 1% (the ticket is polite... see note)

2. News topic โ€” "เธ„เธฃเธก. เน€เธซเน‡เธ™เธŠเธญเธšเธ‚เธถเน‰เธ™เธ„เนˆเธฒเนเธฃเธ‡เธ‚เธฑเน‰เธ™เธ•เนˆเธณเน€เธ›เน‡เธ™ 400 เธšเธฒเธ—เธ—เธฑเนˆเธงเธ›เธฃเธฐเน€เธ—เธจ เธกเธตเธœเธฅ 1 เธกเธเธฃเธฒเธ„เธก เธชเธ เธฒเธญเธธเธ•เธชเธฒเธซเธเธฃเธฃเธกเธเธฑเธ‡เธงเธฅเธเธฃเธฐเธ—เธš SME" (8 topics) โ†’ choice เน€เธจเธฃเธฉเธเธเธดเธˆ 62%, เนเธฃเธ‡เธ‡เธฒเธ™ 36%; noul "เน€เธเธตเนˆเธขเธงเธเธฑเธšเนเธฃเธ‡เธ‡เธฒเธ™เธซเธฃเธทเธญเน„เธกเนˆ" โ†’ 98%. Both readings are right; the probabilities show the overlap instead of hiding it.

3. Assistant intent โ€” "เธŠเนˆเธงเธขเธ•เธฑเน‰เธ‡เธ›เธฅเธธเธเธ•เธญเธ™เธซเธเน‚เธกเธ‡เธ„เธฃเธถเนˆเธ‡เธžเธฃเธธเนˆเธ‡เธ™เธตเน‰เนƒเธซเน‰เธซเธ™เนˆเธญเธข เนเธฅเน‰เธงเธเน‡เน€เธ›เธดเธ”เน€เธžเธฅเธ‡เน€เธšเธฒเน† เธ•เธญเธ™เธ•เธทเนˆเธ™เธ”เน‰เธงเธข" (10 intents) โ†’ alarm set 99.5% (the secondary "play music" request does not distract the primary intent).

4. Comment moderation โ€” "เน„เธญเน‰เธžเธงเธเน€เธซเธตเน‰เธข เธ—เธณเธ‡เธฒเธ™เธเธฑเธ™เนเธšเธšเธ™เธตเน‰เน„เธ›เธ•เธฒเธขเธ‹เธฐ เนƒเธ„เธฃเธเน‡เน„เธ”เน‰เน€เธญเธฒเธ„เธ™เธžเธงเธเธ™เธตเน‰เธญเธญเธเน„เธ›เธ—เธต" โ†’ noul เน€เธ›เน‡เธ™เธžเธดเธฉ 94%; choice sentiment เน€เธŠเธดเธ‡เธฅเธš 97%.

5. Grounded QA (answerable or not) โ€” passage about เธชเธฐเธžเธฒเธ™เธžเธฃเธฐเธฃเธฒเธก 8 (opened 7 May 2545, 475 m long, Bangphlat โ†” Phra Nakhon) โ†’ "เธชเธฐเธžเธฒเธ™เธขเธฒเธงเน€เธ—เนˆเธฒเน„เธฃ" answerable 98.5%; "เธฃเธฐเธšเธธเธ‡เธšเธ›เธฃเธฐเธกเธฒเธ“เธเนˆเธญเธชเธฃเน‰เธฒเธ‡เธซเธฃเธทเธญเน„เธกเนˆ" 0.6%. It says yes only when the fact is actually in the text.

6. Agent tool selection โ€” task "เธˆเธญเธ‡เน‚เธ•เนŠเธฐเธฃเน‰เธฒเธ™เธญเธฒเธซเธฒเธฃ 4 เธ„เธ™ เธ„เธทเธ™เธ™เธตเน‰เธชเธญเธ‡เธ—เธธเนˆเธก", 5 tools, history shows search_restaurants already found a free table โ†’ next tool make_reservation 96.5% (not search again, not SMS yet).

7. RAG relevance + extraction โ€” question on withholding tax for building rent, passage stating 5% โ†’ passage relevant 98%; rate choice 5% (76%), 3% 16%.

Latency from a laptop over the internet was 170โ€“220 ms per request including network; the model itself takes ~40 ms on an H100. Note on example 1's politeness flag: the ticket uses "เธฃเธšเธเธงเธ™โ€ฆเธ™เธฐเธ„เธฃเธฑเธš", so the correct answer is yes; the model said no at 99%. Politeness judgement on formal Thai is a known gap and will get a targeted set in v0.4. We keep the miss here because a showcase that only shows hits is not useful.

Playing Doom, no vision, no text

OpenThai-SystemOne playing Doom

The model controls a Doom marine (ViZDoom) in real time. Every 4 game tics the engine's symbolic state is serialised to text, for example:

health 100/100 | ammo 50 | kills 1
crosshair: empty; nearest visible enemy 39deg to the RIGHT
enemies: Zombieman 5m right -39deg VISIBLE; ChaingunGuy 19m right -24deg; Zombieman 19m ahead -12deg
items: GreenArmor 41m right -18deg
depth ahead: 35/255 (obstacle near)
last actions: ATTACK ATTACK ATTACK ATTACK

and the model answers two typed questions in one forward pass: a choice over the 7 actions (the key that gets pressed) and a noul "is an enemy in the crosshair". About 41 ms per decision on one H100 (~24 decisions/s), 0 output tokens, and no Doom data in training: everything comes from reading the state and the option descriptions. It misses shots and dies on hard levels; the point is the speed and the calibrated probabilities, the same mechanics that route tickets or pick UI elements for an agent.

Run it on your machine (iApp API key or the local weights, live HUD in Thai or English, optional recording): https://github.com/iapp-technology/openthai-systemone-doom

Evaluation

All numbers are zero-shot: the model sees only the state, the instructions and the option names/descriptions.

Public benchmark (same 13 subsets, splits, instructions and sampler as Bespoke Nimble's docs/PUBLIC_BENCHMARKS.md)

Nimble-9B and Jev numbers are as published by Bespoke Labs (2026-09-18); ours are measured with scripts/06_eval.py on scripts/06b_public_benchmarks.py rebuilds of the same subsets.

subset type n OpenThai 0.8B Nimble-9B Jev 1.13.0 our ECE
aegis2 noul 250 83.2 81.2 80.4 0.065
boolq noul 300 79.7 86.0 89.7 0.049
civil_comments noul 300 79.0 70.3 81.0 0.087
helpsteer2 score 250 41.6 39.0 34.1 0.371
massive-de-DE choice 350 88.3 83.4 86.9 0.057
massive-en-US choice 350 88.3 86.9 87.4 0.070
multinli choice 299 89.0 85.3 82.9 0.053
paws noul 250 94.0 82.8 89.2 0.035
pubmedqa choice 250 64.0 75.6 77.2 0.259
squad2 noul 299 89.3 80.6 82.9 0.041
summeval-consistency score 144 75.0 75.7 81.2 0.076
summeval-relevance score 240 21.7 49.2 35.0 0.356
vitaminc-dev choice 599 72.5 76.6 80.1 0.118
macro average 74.3 74.8 76.0

For scale: Bespoke reports raw Qwen3.5-0.8B at 45.4 on their private 324-item holdout (not this bench), Nimble-9B at 90.1, Jev at 93.2. choice/noul report accuracy; score reports exact-level match. Bold = ahead of Bespoke-Nimble-9B. Honest reading: the 0.8B model is ahead of the 9B on 4 of 13 subsets (NLI, summary consistency, helpfulness scoring, toxicity) and clearly behind on reading-comprehension style yes/no tasks (squad2 is at chance, boolq, pubmedqa) and on summary relevance scoring, which is the one subset where our score head is badly miscalibrated (ECE 0.79).

Thai held-out sets (never in training; whole datasets held out where marked)

set type n accuracy macro-F1 / MAE ECE note
MASSIVE-th intent (60-way) choice 5007 90.0 F1 0.869 0.043 eval split
Prachathai67k topics choice 3501 98.1 F1 0.938 0.004 eval split
Prachathai67k topics noul 13119 94.2 0.008 eval split
XNLI-th choice 2490 77.1 F1 0.772 0.045 eval split
XNLI-th (entailment yes/no) noul 2490 84.3 0.049 eval split
SIB-200 Thai topic (7-way) choice 204 77.9 F1 0.759 0.084 whole dataset held out (v0.1: 77.5)
Thai contrastive pairs (one-fact flips) choice 296 80.7 F1 0.734 0.098 synthetic, eval-only
Thai contrastive pairs score 56 78.6 MAE 0.35 0.156 synthetic, eval-only
Thai contrastive pairs noul 248 83.5 0.109 synthetic, eval-only
Wongnai review stars (1โ€“5) score 6203 63.5 MAE 0.44 0.039 eval split
Wisesight sentiment (4-class) choice 2671 51.6 F1 0.448 0.353 whole dataset held out โ€” v0.1 38.7 โ†’ v0.2 51.5 โ†’ now 51.6 (weakest Thai set; use order-invariant mode)
banking77 intent (77-way, English) choice 3076 45.4 F1 0.417 0.236 whole dataset held out โ€” 77-way near-duplicate intents (v0.1 32.7; 61.7 with order-invariant mode on v0.2)
xLAM tool selection (English) choice 884 99.4 F1 0.986 0.006 eval slice

Batch-1 latency, one question with 255 options, H100 shared with a training job: 44 ms (public bench run), 48 ms (held-out run). A 3-question Thai ticket (166 tokens): ~40 ms on H100, 154 ms on a MacBook M3 Max (MPS).

Through Ollama (up to 26 options per question)

All models below were run by us through Ollama 0.35's /v1/systemone on the same records: the first 800 of each set, keeping only records whose questions have โ‰ค 26 options, since Ollama rejects more. That drops banking77 and the 60-way MASSIVE-th intent questions. "This repo's API" is the v0.3 weights with the slot head on the same records.

model size public 13 (macro) Thai 8 sets / 12 rows (macro) latency, 1 / 3 questions
OpenThai-SystemOne v0.3, this repo's API 0.8B 74.3 82.2 โ€“
OpenThai-SystemOne v0.3, Ollama build Q8_0 0.8B 74.3 82.5 23 / 136 ms
Ollama build Q4_K_M 0.8B 74.0 81.6 23 / 127 ms
Tev1 0.8B (Together AI, tev1:0.8b) 0.8B 64.3 62.7 22 / 102 ms
Tev1 4B (tev1:4b) 4B 74.9 75.4 66 / 406 ms
Nimble 9B (Bespoke Labs, nimble) 9B 74.8 78.3 69 / 392 ms

Latency: median end to end through Ollama, one model loaded, idle H100, Thai requests. Ollama runs one prompt per question. The Ollama build ties this repo's API. It leads the other 0.8B model by 10 points (public) and 20 points (Thai), and on Thai it also leads the 4B and 9B models (82.5 vs 75.4 / 78.3). On English the 4B and 9B are within a point of it. They are ahead on reading tasks (BoolQ, PubMedQA, VitaminC, SummEval-relevance) and on the Thai contrastive pairs. The Ollama build is ahead on Thai topic, rating and NLI (Prachathai 97.8 vs 61.5, Wongnai 64.0 vs 52.1, XNLI-th 79.2 vs 76.0 for Nimble), on PAWS and SQuAD2. Per-set tables: iapp/OpenThai-SystemOne-Ollama.

Calibration (Stage 3)

v0.3 learned temperatures: choice 1.055, noul 1.047, score 1.008 (the before/after table below was measured on v0.1; the procedure is identical in every version).

After SFT (12k steps) the backbone was frozen and the slot head plus one temperature per question type were trained for 400 steps on the SFT mixture with cross-entropy + Brier loss (configs/calib.yaml, brier_weight: 1.0, train_temperature: true). Learned temperatures: choice 1.062, noul 1.047, score 1.008.

Before/after on the same sources (before = SFT checkpoint, 150โ€“300-item smoke slices; after = calibrated checkpoint, full slices, so accuracies are not strictly comparable, ECE is):

set type ECE before โ†’ after accuracy before โ†’ after
multinli choice 0.049 โ†’ 0.035 84.7 โ†’ 85.6
massive-en-US choice 0.064 โ†’ 0.133 73.3 โ†’ 75.7
boolq noul 0.086 โ†’ 0.195 69.3 โ†’ 63.7
paws noul 0.108 โ†’ 0.143 69.3 โ†’ 67.2
vitaminc-dev choice 0.104 โ†’ 0.098 58.0 โ†’ 67.1
MASSIVE-th choice 0.032 โ†’ 0.048 89.7 โ†’ 86.4
Prachathai (choice / noul) 0.033 / 0.024 โ†’ 0.005 / 0.008 97.4 / 93.4 โ†’ 97.7 / 94.1
XNLI-th (choice / noul) 0.062 / 0.012 โ†’ 0.028 / 0.042 80.0 / 87.0 โ†’ 76.5 / 84.3
SIB-200 th choice 0.062 โ†’ 0.074 77.5 โ†’ 77.5
Wongnai stars score 0.071 โ†’ 0.010 66.3 โ†’ 63.3
Thai contrastive (choice / noul) 0.083 / 0.075 โ†’ 0.105 / 0.088 79.1 / 84.6 โ†’ 78.7 / 82.3
Wisesight choice 0.286 โ†’ 0.341 39.0 โ†’ 38.7

Reading: calibration helps where the model is already competent (Thai topic/NLI, Wongnai scoring, MultiNLI: ECE โ‰ค 0.05) and does not rescue sets where accuracy itself is low (wisesight, squad2, summeval-relevance) โ€” a temperature cannot fix a wrong ranking. On the English public bench the median ECE is 0.15; treat confidence as reliable on the Thai sets and on NLI/topic tasks, and route low-confidence English yes/no decisions to a bigger model.

Limits

  • Text only. Up to 255 options per question in one stage (bucket into groups for more). 64k tokens per request.
  • It is a small model: use the confidence field and route low-confidence cases to a bigger model or a human.
  • Not a reasoning model: it will not do multi-step verification or arithmetic.
  • v0.3 known weak spots (numbers above): summary relevance scoring (22) and helpfulness scoring (42) โ€” the two 5-level rating tasks; fine-grained 77-way English intents (banking77 45 single-order); Thai social sentiment (wisesight 52); PubMedQA 3-way (64).
  • English calibration is weaker than Thai (median ECE 0.15 vs โ‰ค 0.05): the training mix is Thai-heavy by design.
  • Ollama build: โ‰ค 26 options per question (Ollama's limit), no order-invariant mode, no abstain; use this repo's API for more.

Versions

version date change public macro Wisesight
v0.3 Ollama build 2026-09-30 v0.3 fine-tuned for Ollama 0.35's /v1/systemone prompt: ollama.com/iapp/openthai-systemone, GGUF Q8_0 / Q4_K_M / BF16 in iapp/OpenThai-SystemOne-Ollama 74.3 (โ‰ค 26 options) 49.9 (first 800)
v0.3 2026-09-22 +5,000 SFT steps from v0.2 with 177k real train-split records + 78k targeted synthetic records for the weak spots (grounded QA, summary rating, fine-grained intents, safety, paraphrase), re-calibrated 74.3 51.6
v0.2 2026-09-21 +3,000 SFT steps from v0.1 with a 22k-record synthetic Thai social-sentiment set (4/3/5-class, yes/no, score schemes), re-calibrated 63.2 51.5
v0.1 2026-09-20 initial release: Thai CPT 4.47B tokens, 12k-step SFT, calibration 61.9 38.7

Changelog

v0.3 Ollama build โ€” 2026-09-30

  • Ollama build for Ollama โ‰ฅ 0.35's System One API: ollama pull iapp/openthai-systemone (ollama.com) or ollama pull hf.co/iapp/OpenThai-SystemOne-Ollama:Q8_0 (GGUF repo), then POST /v1/systemone. The weights in this repo are unchanged.
  • How: Ollama renders each question as a Qwen3.5 chat prompt and scores the answer letters, so v0.3 was fine-tuned for 3,000 steps (192k questions, same data mix as v0.3) on prompts built by a byte-exact port of Ollama's code, tokenized by llama.cpp as Ollama does, with cross-entropy over the candidate letters and one fitted temperature (0.83) folded into the weights.
  • Through Ollama: public 13 macro 74.3, Thai 82.5 (this repo's API on the same records: 74.3 / 82.2); Q4_K_M 74.0 / 81.6.

v0.3 โ€” 2026-09-22

  • Continued fine-tuning for 5,000 steps from v0.2 with the weak-spot data: real train splits (SQuAD2, BoolQ, PubMedQA-artificial, PAWS, Aegis2, ToxicChat, XQuAD-th, MASSIVE-de; 177k records) and five targeted synthetic sets generated with Qwen3.6-35B-A3B and blind-checked (grounded yes/no QA with near-miss unanswerables 30k, summary rating on the SummEval rubrics 20k, 40โ€“120-way near-duplicate intent taxonomies 12k, safety judgments 8k, adversarial paraphrases 8k); sentiment set weight lowered from 4ร— to 2ร—; re-calibrated.
  • Public 13-subset macro 63.2 โ†’ 74.3 (Nimble-9B 74.8): SQuAD2 50.2 โ†’ 89.3, PAWS 68.0 โ†’ 94.0, Aegis2 61.6 โ†’ 83.2, MASSIVE-de 67.4 โ†’ 88.3, BoolQ 64.7 โ†’ 79.7, MASSIVE-en 79.1 โ†’ 88.3, PubMedQA 56.4 โ†’ 64.0, SummEval-relevance 14.2 โ†’ 21.7, VitaminC 68.6 โ†’ 72.5, MultiNLI 87.3 โ†’ 89.0. ECE improved on 11 of 13 subsets.
  • Regression: SummEval-consistency 84.0 โ†’ 75.0 (the synthetic consistency set only covers levels 1/3/5; levels 2/4 will be added). Thai held-out sets all flat or up (SIB-200 74.0 โ†’ 77.9, banking77 37.8 โ†’ 45.4 single-order).
  • Note: SQuAD2, BoolQ, PAWS, Aegis2 and MASSIVE-de are no longer "never trained on": their train splits are now in the mix; the public-bench subsets still use the validation/test splits only.

v0.2 โ€” 2026-09-21

  • Continued fine-tuning for 3,000 steps from v0.1 with a new 22k-record synthetic Thai social-media sentiment set (4-class with the "question" class, 3/5-class variants, yes/no flags, 5-level score), then re-calibrated.
  • Public 13-subset macro 61.9 โ†’ 63.2; Wisesight 38.7 โ†’ 51.5; MASSIVE-en 75.7 โ†’ 79.1; MASSIVE-th 86.4 โ†’ 88.6; MultiNLI 85.6 โ†’ 87.3 (ECE 0.035 โ†’ 0.016); Aegis2 58.0 โ†’ 61.6; PubMedQA 53.6 โ†’ 56.4.
  • Regression: SIB-200 Thai 77.5 โ†’ 74.0 (sentiment set was weighted 4ร—; will be lowered next round). Unchanged: SQuAD2, SummEval-relevance.
  • Known issue and fix: like Jev, a single forward pass is sensitive to the order of the options (measured on v0.2: 5โ€“12% of arg-max choices flip under a random permutation for โ‰ค 10 options, 72% for 77-way banking77). The client and API now have an order-invariant mode that averages the answer over several cyclic option orders in one batched pass (order_invariant: true / permutations: n; automatic for choice questions with > 10 options). With it, on 360 held-out Thai questions: flips 18.9% โ†’ 3.3%, mean probability shift 0.19 โ†’ 0.08, accuracy 68.6 โ†’ 73.6, banking77 41.7 โ†’ 63.3. Cost: ~2ร— latency (batched), never more tokens per request in usage.input_tokens.
  • Live API switched to v0.2 (same checkpoint as these weights).

v0.1 โ€” 2026-09-20

  • Initial release: Qwen3.5-0.8B text tower, 4.47B-token Thai CPT, 256-slot decision head, 12k-step SFT on 1.8M public + 98k synthetic decision records, calibration stage. Public macro 61.9.

License & credits

Apache-2.0. Built by iApp Technology / OpenThai on Qwen3.5-0.8B-Base (Apache-2.0). Inspired by TypeSafe AI's System One models (Jev); this is an independent open re-implementation and is not affiliated with TypeSafe AI.

Sponsor

Siam AI Corporation

Training, evaluation and the free hosted API for this model run on NVIDIA H100 GPUs generously provided by Siam AI Corporation. Thank you for backing open Thai AI.

Downloads last month
9,737
Safetensors
Model size
0.8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for iapp/OpenThai-SystemOne

Finetuned
(124)
this model
Finetunes
3 models
Quantizations
14 models

Spaces using iapp/OpenThai-SystemOne 2

Collection including iapp/OpenThai-SystemOne