📊 Thai LLM Leaderboard

ผลทดสอบโมเดลภาษาไทยภายใต้มาตรฐานเดียวกัน — Thai language models, one identical protocol.

How to read this. Scores are only comparable within a section: each section is one identical protocol — same prompts, same harness, every model re-run by us — and the boldface best-in-column is computed, not chosen. The leaderboard is maintained by OpenThai and includes models that beat ours on individual benchmarks. OpenThaiEval is our own exam benchmark; treat that column accordingly.

Protocol B (September 2026) — the OpenThai 2.0 release evaluation: EvalScope for OpenThaiEval / MMLU-Redux (temperature 0), the official IFEval and bfcl-eval harnesses, CER against human-verified ground truth for document reading, and the NitiBench citation-F1 scorer for law. Single pass, 16k context; thinking on for text tasks and off for image-attached OCR and for citation answering. Scores are 0–1. Full protocol on the OpenThai 2.0 and OpenThai 2.0 Legal model cards.

Protocol A (April 2025) — one SkyThought evaluation run from the OpenThaiGPT 1.6 & R1 technical report. Scores are 0–100; -TH marks Thai-language variants. Not re-run for the 2.0 generation, so do not read across the two protocols.

Protocol B · Thai knowledge, instruction following and code — ↑ higher is better

OpenThai 2.0 generation, September 2026. 0–1, higher is better. Average is computed over the six benchmarks.

Protocol B · Thai document reading (OCR) — character error rate, ↓ lower is better

Lower is better in every column of this table (CER against human-verified ground truth, 0 = perfect). Typhoon-OCR 1.5 is a transcription-only model (fixed OCR prompt, no assistant ability) and wins the clean printed and handwriting sets; OpenThai 2.0 is the only model in the table that both reads documents and answers questions about them.

Protocol B · Agentic tool use (BFCL) — ↑ higher is better

Berkeley Function-Calling Leaderboard, bfcl-eval 2026.3.23: 3,841 cases in 14 categories, identical OpenAI function-calling protocol. 0–1, higher is better.

Protocol B · Thai law (OpenThai 2.0 Legal) — ↑ higher is better

Closed-book = recall the statute from memory; open-book = cite from provided text (the RAG setting); essays = 72 held-out Supreme Court cases, citations checked by code, holding / coverage / fluency judged by Gemini 3.1 Flash Lite. Single pass, thinking off. 0–1, higher is better.

Protocol A · Reasoning models

Step-by-step reasoning models, evaluated with thinking enabled.

Protocol A · General instruction models

General-purpose chat models. Language Accuracy = replies in the language it was asked in.