Benchmark results

Lanna LLM Eval Leaderboard

Model performance across Lanna-language understanding benchmarks. The current snapshot uses 31 Lanna-translated questions from Global MMLU Lite, ranked by overall accuracy.

Models evaluated

17

Questions per model

31

Top accuracy

93.55%

Accuracy comparison

Overall accuracy across all 31 questions.

1 Claude Opus 5

93.55%

2 DeepSeek V4 Pro 0813

93.55%

3 Gemini 3.1 Pro Preview

93.55%

4 GPT-5.6 Sol

93.55%

5 Gemini 3.7 Flash

93.10%

6 GPT-5.6 Terra

87.10%

7 Claude Sonnet 5

83.87%

8 GPT-5.6 Luna

83.87%

9 Grok 4.6

83.87%

10 Qwen 3.8 2.4T A95B

70.97%

11 DeepSeek V4 Flash 0731

58.06%

12 Kimi K3

58.06%

13 Qwen 3.8 27B

54.84%

14 Qwen 3.8 Max

51.61%

15 HY3

48.39%

16 MiniMax M3

38.71%

17 GLM 5.2

35.48%

Overall ranking

Higher accuracy is better.

Rank Model Method Correct Accuracy Macro F1
1
Claude Opus 5
Anthropic
Manual 29/31 93.55% 93.98%
2
DeepSeek V4 Pro 0813
DeepSeek
Max 29/31 93.55% 93.05%
3
Gemini 3.1 Pro Preview
Google
Max 29/31 93.55% 93.05%
4
GPT-5.6 Sol
OpenAI
Manual 29/31 93.55% 93.05%
5
Gemini 3.7 Flash
Google
Max 27/31 93.10% 92.44%
6
GPT-5.6 Terra
OpenAI
Manual 27/31 87.10% 88.16%
7
Claude Sonnet 5
Anthropic
Manual 26/31 83.87% 85.25%
8
GPT-5.6 Luna
OpenAI
Manual 26/31 83.87% 81.76%
9
Grok 4.6
xAI
Max 26/31 83.87% 83.01%
10
Qwen 3.8 2.4T A95B
Qwen
Max 22/31 70.97% 70.18%
11
DeepSeek V4 Flash 0731
DeepSeek
Max 18/31 58.06% 58.52%
12
Kimi K3
Moonshot AI
Max 18/31 58.06% 57.46%
13
Qwen 3.8 27B
Qwen
Medium · hint 17/31 54.84% 55.30%
14
Qwen 3.8 Max
Qwen
Max 16/31 51.61% 49.27%
15
HY3
Tencent
Max 15/31 48.39% 49.35%
16
MiniMax M3
MiniMax
Max 12/31 38.71% 40.48%
17
GLM 5.2
Z.ai
Max 11/31 35.48% 35.40%

Accuracy by category

Percentage of correct answers in each benchmark category.

Model Business Humanities Medical Other STEM Social Sciences
Claude Opus 5 100.00% 85.71% 100.00% 100.00% 100.00% 75.00%
DeepSeek V4 Pro 0813 100.00% 85.71% 100.00% 100.00% 100.00% 75.00%
Gemini 3.1 Pro Preview 100.00% 85.71% 100.00% 100.00% 100.00% 75.00%
GPT-5.6 Sol 100.00% 85.71% 100.00% 100.00% 100.00% 75.00%
Gemini 3.7 Flash 100.00% 85.71% 100.00% 100.00% 100.00% 66.67%
GPT-5.6 Terra 75.00% 85.71% 100.00% 100.00% 100.00% 75.00%
Claude Sonnet 5 100.00% 71.43% 75.00% 75.00% 75.00% 100.00%
GPT-5.6 Luna 87.50% 57.14% 100.00% 100.00% 100.00% 75.00%
Grok 4.6 100.00% 71.43% 100.00% 75.00% 100.00% 50.00%
Qwen 3.8 2.4T A95B 100.00% 57.14% 100.00% 75.00% 75.00% 0.00%
DeepSeek V4 Flash 0731 50.00% 42.86% 50.00% 75.00% 100.00% 50.00%
Kimi K3 87.50% 14.29% 75.00% 50.00% 50.00% 75.00%
Qwen 3.8 27B 75.00% 42.86% 50.00% 75.00% 50.00% 25.00%
Qwen 3.8 Max 75.00% 42.86% 50.00% 75.00% 50.00% 0.00%
HY3 87.50% 14.29% 75.00% 25.00% 50.00% 25.00%
MiniMax M3 75.00% 14.29% 50.00% 0.00% 50.00% 25.00%
GLM 5.2 50.00% 42.86% 25.00% 50.00% 0.00% 25.00%

This snapshot was scored on 15 August 2026. Each model answered the same 31 Lanna-translated questions across Business, Humanities, Medical, Other, STEM, and Social Sciences. “Manual” results were scored from supplied answers; “Max” used the evaluation run's maximum reasoning effort.

Read the accompanying research notes on Lanna and LLMs and the Global MMLU Lite dataset .