Benchmark results
Lanna LLM Eval Leaderboard
Model performance across Lanna-language understanding benchmarks. The current snapshot uses 31 Lanna-translated questions from Global MMLU Lite, ranked by overall accuracy.
Models evaluated
17
Questions per model
31
Top accuracy
93.55%
Accuracy comparison
Overall accuracy across all 31 questions.
1 Claude Opus 5
93.55%
2 DeepSeek V4 Pro 0813
93.55%
3 Gemini 3.1 Pro Preview
93.55%
4 GPT-5.6 Sol
93.55%
5 Gemini 3.7 Flash
93.10%
6 GPT-5.6 Terra
87.10%
7 Claude Sonnet 5
83.87%
8 GPT-5.6 Luna
83.87%
9 Grok 4.6
83.87%
10 Qwen 3.8 2.4T A95B
70.97%
11 DeepSeek V4 Flash 0731
58.06%
12 Kimi K3
58.06%
13 Qwen 3.8 27B
54.84%
14 Qwen 3.8 Max
51.61%
15 HY3
48.39%
16 MiniMax M3
38.71%
17 GLM 5.2
35.48%
Overall ranking
Higher accuracy is better.
| Rank | Model | Method | Correct | Accuracy | Macro F1 |
|---|---|---|---|---|---|
| 1 | Claude Opus 5 Anthropic | Manual | 29/31 | 93.55% | 93.98% |
| 2 | DeepSeek V4 Pro 0813 DeepSeek | Max | 29/31 | 93.55% | 93.05% |
| 3 | Gemini 3.1 Pro Preview Google | Max | 29/31 | 93.55% | 93.05% |
| 4 | GPT-5.6 Sol OpenAI | Manual | 29/31 | 93.55% | 93.05% |
| 5 | Gemini 3.7 Flash Google | Max | 27/31 | 93.10% | 92.44% |
| 6 | GPT-5.6 Terra OpenAI | Manual | 27/31 | 87.10% | 88.16% |
| 7 | Claude Sonnet 5 Anthropic | Manual | 26/31 | 83.87% | 85.25% |
| 8 | GPT-5.6 Luna OpenAI | Manual | 26/31 | 83.87% | 81.76% |
| 9 | Grok 4.6 xAI | Max | 26/31 | 83.87% | 83.01% |
| 10 | Qwen 3.8 2.4T A95B Qwen | Max | 22/31 | 70.97% | 70.18% |
| 11 | DeepSeek V4 Flash 0731 DeepSeek | Max | 18/31 | 58.06% | 58.52% |
| 12 | Kimi K3 Moonshot AI | Max | 18/31 | 58.06% | 57.46% |
| 13 | Qwen 3.8 27B Qwen | Medium · hint | 17/31 | 54.84% | 55.30% |
| 14 | Qwen 3.8 Max Qwen | Max | 16/31 | 51.61% | 49.27% |
| 15 | HY3 Tencent | Max | 15/31 | 48.39% | 49.35% |
| 16 | MiniMax M3 MiniMax | Max | 12/31 | 38.71% | 40.48% |
| 17 | GLM 5.2 Z.ai | Max | 11/31 | 35.48% | 35.40% |
Accuracy by category
Percentage of correct answers in each benchmark category.
| Model | Business | Humanities | Medical | Other | STEM | Social Sciences |
|---|---|---|---|---|---|---|
| Claude Opus 5 | 100.00% | 85.71% | 100.00% | 100.00% | 100.00% | 75.00% |
| DeepSeek V4 Pro 0813 | 100.00% | 85.71% | 100.00% | 100.00% | 100.00% | 75.00% |
| Gemini 3.1 Pro Preview | 100.00% | 85.71% | 100.00% | 100.00% | 100.00% | 75.00% |
| GPT-5.6 Sol | 100.00% | 85.71% | 100.00% | 100.00% | 100.00% | 75.00% |
| Gemini 3.7 Flash | 100.00% | 85.71% | 100.00% | 100.00% | 100.00% | 66.67% |
| GPT-5.6 Terra | 75.00% | 85.71% | 100.00% | 100.00% | 100.00% | 75.00% |
| Claude Sonnet 5 | 100.00% | 71.43% | 75.00% | 75.00% | 75.00% | 100.00% |
| GPT-5.6 Luna | 87.50% | 57.14% | 100.00% | 100.00% | 100.00% | 75.00% |
| Grok 4.6 | 100.00% | 71.43% | 100.00% | 75.00% | 100.00% | 50.00% |
| Qwen 3.8 2.4T A95B | 100.00% | 57.14% | 100.00% | 75.00% | 75.00% | 0.00% |
| DeepSeek V4 Flash 0731 | 50.00% | 42.86% | 50.00% | 75.00% | 100.00% | 50.00% |
| Kimi K3 | 87.50% | 14.29% | 75.00% | 50.00% | 50.00% | 75.00% |
| Qwen 3.8 27B | 75.00% | 42.86% | 50.00% | 75.00% | 50.00% | 25.00% |
| Qwen 3.8 Max | 75.00% | 42.86% | 50.00% | 75.00% | 50.00% | 0.00% |
| HY3 | 87.50% | 14.29% | 75.00% | 25.00% | 50.00% | 25.00% |
| MiniMax M3 | 75.00% | 14.29% | 50.00% | 0.00% | 50.00% | 25.00% |
| GLM 5.2 | 50.00% | 42.86% | 25.00% | 50.00% | 0.00% | 25.00% |
This snapshot was scored on 15 August 2026. Each model answered the same 31 Lanna-translated questions across Business, Humanities, Medical, Other, STEM, and Social Sciences. “Manual” results were scored from supplied answers; “Max” used the evaluation run's maximum reasoning effort.
Read the accompanying research notes on Lanna and LLMs and the Global MMLU Lite dataset .