How to read this. Every score below comes from a
single evaluation run
on the SkyThought harness, published in the
OpenThaiGPT 1.6 & R1 technical report
(April 2025) — same prompts, same protocol, every model. Scores are 0–100, higher is better;
-TH marks Thai-language variants. The leaderboard is maintained by OpenThai and includes
models that beat ours on individual benchmarks — the boldface best-in-column is computed, not chosen.
OpenThaiEval is our own exam benchmark; treat that column accordingly.
Step-by-step reasoning models, evaluated with thinking enabled.
General-purpose chat models. Language Accuracy = replies in the language it was asked in.