How to read this. Scores are only comparable
within a section: each
section is one identical protocol — same prompts, same harness, every model re-run by us — and the
boldface best-in-column is computed, not chosen. The leaderboard is maintained by OpenThai and
includes models that beat ours on individual benchmarks. OpenThaiEval is our own exam benchmark;
treat that column accordingly.
Protocol B (September 2026) — the OpenThai 2.0 release evaluation: EvalScope for
OpenThaiEval / MMLU-Redux (temperature 0), the official IFEval and
bfcl-eval harnesses,
CER against human-verified ground truth for document reading, and the
NitiBench
citation-F1 scorer for law. Single pass, 16k context; thinking on for text tasks and off for
image-attached OCR and for citation answering. Scores are 0–1. Full protocol on the
OpenThai 2.0 and
OpenThai 2.0 Legal model cards.
Protocol A (April 2025) — one SkyThought evaluation run from the
OpenThaiGPT 1.6 & R1 technical report.
Scores are 0–100;
-TH marks Thai-language variants. Not re-run for the 2.0 generation, so do not
read across the two protocols.
Protocol B · Thai knowledge, instruction following and code — ↑ higher is better
OpenThai 2.0 generation, September 2026. 0–1, higher is better. Average is computed over the six benchmarks.
Protocol B · Thai document reading (OCR) — character error rate, ↓ lower is better
Lower is better in every column of this table (CER against human-verified ground truth, 0 = perfect). Typhoon-OCR 1.5 is a transcription-only model (fixed OCR prompt, no assistant ability) and wins the clean printed and handwriting sets; OpenThai 2.0 is the only model in the table that both reads documents and answers questions about them.
Protocol B · Agentic tool use (BFCL) — ↑ higher is better
Berkeley Function-Calling Leaderboard, bfcl-eval 2026.3.23: 3,841 cases in 14 categories, identical OpenAI function-calling protocol. 0–1, higher is better.
Protocol B · Thai law (OpenThai 2.0 Legal) — ↑ higher is better
Closed-book = recall the statute from memory; open-book = cite from provided text (the RAG setting); essays = 72 held-out Supreme Court cases, citations checked by code, holding / coverage / fluency judged by Gemini 3.1 Flash Lite. Single pass, thinking off. 0–1, higher is better.
Protocol A · Reasoning models
Step-by-step reasoning models, evaluated with thinking enabled.
Protocol A · General instruction models
General-purpose chat models. Language Accuracy = replies in the language it was asked in.