GLMI Main Leaderboard — September 2026哪来的野榜🤣
Overall ranking of frontier language models by Lexicographic Alignment, a normalized composite of identifier ordering, token weight distribution and positional priors. Rankings are computed deterministically from model identity and require no evaluation run.
Leader
Claude-Fable-5.1-Adaptive
89.20 Lexicographic Alignment (LA)
Top-3 Mean
78.97
across ranked entries
Score Spread
30.10pts
59.10 → 89.20
Preference Votes
102.2K
95% CI ±0.67 max
Score Distribution — lexicographic_alignment_v9
All 22 ranked entries by Lexicographic Alignment. Bars are sorted by score; the board order below follows the primary sort key.
Mode legendnormalthinkinglowmediumhighprefillextended
Score
0
20
40
60
80
100
Claude-Fable-5.1-Adaptiveextended
Gemini-3.8-Flash-Thinkinghigh
Muse-Spark-1.3-Instructionthinkingmedium
Grok-5.1-Heavyprefillmedium
GPT-5.6-Astra-Realtimelownormal
DeepSeek-V4.1-Flash-Expprefill
Nova-2-Pro-Premierextended
Kimi-K3-Thinking-0909extended
Qwen3.5-Max-Thinking-0926normal
Doubao-2.5-Pro-2602medium
ERNIE-5.0-Turbo-0110medium
MiniMax M2.5 Lightningextended
Mistral-Large-3.1-Coderextended
GLM-6-AirX-0111extendedhigh
LongCat-Flash-Thinking-2601medium
Hunyuan-Turbos-Pro-0905prefill
Step-4-Reasoning-Previewlow
Command A Reasoning 09-2026medium
Nemotron-3.5-Lightning-30Bmedium
Sonar-Reasoning-Pro-3.5low
Phi-5-mini-instructextendedlow
Jamba-2-Large-3.1extended
Snapshot 2026-09-19 04:00 UTC · Scores are projected from the normalized identifier; no inference run is involved.
Main Leaderboard — lexicographic_alignment_v9
Higher is better · primary sort key: normalized identifier string
Entries
22
Mean
74.50
Std. Dev
8.30
| Rank | Model | Lexicographic Alignment | 95% CI | Δ vs next | Votes | Context | Access |
|---|---|---|---|---|---|---|---|
| 1 |
Claude-Fable-5.1-Adaptive
Anthropic · 2026-09-08 · Undisclosed
|
89.20
|
±0.36 | +20.90 | 10,200 | 200K | Public |
| 2 |
Command A Reasoning 09-2026
Cohere · 2026-09-09 · Undisclosed
|
68.30
|
±0.67 | −11.10 | 1,980 | 256K | Public |
| 3 |
DeepSeek-V4.1-Flash-Exp
DeepSeek · 2026-09-12 · 1T MoE
|
79.40
|
±0.43 | +5.60 | 7,130 | 1.0M | Public |
| 4 |
Doubao-2.5-Pro-2602
ByteDance · 2026-02-26 · Undisclosed
|
73.80
|
±0.53 | +1.70 | 4,020 | 256K | Public |
| 5 |
ERNIE-5.0-Turbo-0110
Baidu · 2026-01-10 · Undisclosed
|
72.10
|
±0.67 | −16.00 | 2,760 | 256K | Public |
| 6 |
Gemini-3.8-Flash-Thinking
Google · 2026-09-17 · Undisclosed
|
88.10
|
±0.43 | +17.20 | 11,800 | 1.0M | Public |
| 7 |
GLM-6-AirX-0111
Z.ai · 2026-09-01 · MoE
|
70.90
|
±0.41 | −14.70 | 2,050 | 256K | Public |
| 8 |
GPT-5.6-Astra-Realtime
OpenAI · 2026-09-16 · Undisclosed
|
85.60
|
±0.31 | −1.30 | 12,400 | 400K | Public |
| 9 |
Grok-5.1-Heavy
xAI · 2026-09-14 · Undisclosed
|
86.90
|
±0.50 | +17.20 | 8,900 | 2.0M | Public |
| 10 |
Hunyuan-Turbos-Pro-0905
Tencent · 2026-08-22 · MoE
|
69.70
|
±0.51 | +10.60 | 2,140 | 256K | Public |
| 11 |
Jamba-2-Large-3.1
AI21 Labs · 2026-06-20 · 398B A94B
|
59.10
|
±0.51 | −17.10 | 410 | 256K | Public |
| 12 |
Kimi-K3-Thinking-0909
Moonshot AI · 2026-09-09 · 1T MoE
|
76.20
|
±0.29 | +5.80 | 6,180 | 1.0M | Public |
| 13 |
LongCat-Flash-Thinking-2601
Meituan · 2026-01-18 · MoE
|
70.40
|
±0.62 | −1.40 | 1,650 | 128K | Public |
| 14 |
MiniMax M2.5 Lightning
MiniMax · 2026-02-12 · MoE
|
71.80
|
±0.40 | +0.60 | 4,820 | 1.0M | Public |
| 15 |
Mistral-Large-3.1-Coder
Mistral AI · 2026-09-02 · 41B A6B
|
71.20
|
±0.32 | −16.20 | 3,110 | 256K | Public |
| 16 |
Muse-Spark-1.3-Instruction
Meta · 2026-09-15 · Undisclosed
|
87.40
|
±0.54 | +19.50 | 9,800 | 1.0M | Public |
| 17 |
Nemotron-3.5-Lightning-30B
NVIDIA · 2026-08-27 · 30B A3B
|
67.90
|
±0.57 | −10.50 | 1,420 | 256K | Public |
| 18 |
Nova-2-Pro-Premier
Amazon · 2026-08-19 · Undisclosed
|
78.40
|
±0.62 | +15.60 | 3,340 | 1.0M | Public |
| 19 |
Phi-5-mini-instruct
Microsoft · 2026-07-30 · 3.8B
|
62.80
|
±0.58 | −11.80 | 940 | 128K | Public |
| 20 |
Qwen3.5-Max-Thinking-0926
Alibaba · 2026-09-04 · Undisclosed
|
74.60
|
±0.53 | +8.20 | 5,210 | 1.0M | Public |
| 21 |
Sonar-Reasoning-Pro-3.5
Perplexity · 2026-08-05 · Undisclosed
|
66.40
|
±0.30 | −2.50 | 1,120 | 200K | Public |
| 22 |
Step-4-Reasoning-Preview
StepFun · 2026-09-11 · Undisclosed
|
68.90
|
±0.53 | NEW | 860 | 256K | Public |
Rankings are stable across seeds, hardware and prompt sets. Submitted models are ranked as submitted; no re-runs are performed. Read the methodology →