Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
29.3
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 101 | Qwen3-Next-80B-A3B-Thinking qwen3-next-80b-a3b-thinking textinference | Alibaba Cloud / Qwen Team | 42.0 overall | 42.3 | 0.0 | 41.7 | 0.0 | 0.0 | N/A |
| 102 | Qwen3 VL 235B A22B Instruct qwen3-vl-235b-a22b-instruct multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 42.0 overall | 34.3 | 0.0 | 51.0 | 0.0 | 0.0 | |
| 103 | Qwen3-Coder 480B A35B Instruct qwen3-coder-480b-a35b-instruct codeprogrammingtool use | Alibaba Cloud / Qwen Team | 41.8 overall | 0.0 | 0.0 | 50.7 | 32.1 | 0.0 | |
| 104 | Qwen3.5-35B-A3B qwen3.5-35b-a3b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 41.6 overall | 53.4 | 0.0 | 38.3 | 30.2 | 0.0 | N/A |
| 105 | GLM-4.6 glm-4.6 multimodalvisionmulti-input reasoning | Zhipu AI | 41.2 overall | 44.5 | 0.0 | 35.4 | 43.2 | 0.0 | N/A |
| 106 | Kimi K2 0905 kimi-k2-0905 textinference | Moonshot AI | 41.0 overall | 41.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 107 | Qwen3-235B-A22B-Instruct-2507 qwen3-235b-a22b-instruct-2507 textinference | Alibaba Cloud / Qwen Team | 40.7 overall | 40.7 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 108 | LongCat-Flash-Lite longcat-flash-lite codeprogrammingtool use | Meituan | 40.1 overall | 22.8 | 72.8 | 30.1 | 23.9 | 95.6 | $0.1 in / $0.4 out |
| 109 | Gemini 2.5 Pro Preview 06-05 gemini-2.5-pro-preview-06-05 multimodalvisionmulti-input reasoning | Google | 39.4 overall | 50.2 | 0.0 | 0.0 | 25.8 | 0.0 | |
| 110 | DeepSeek-V3.2-Exp deepseek-v3.2-exp codeprogrammingtool use | DeepSeek | 39.1 overall | 50.2 | 0.0 | 27.2 | 38.0 | 0.0 | N/A |
| 111 | LongCat-Flash-Thinking-2601 longcat-flash-thinking-2601 codeprogrammingtool use | Meituan | 38.3 overall | 52.9 | 0.0 | 25.6 | 33.5 | 0.0 | |
| 112 | Mistral Small 4 mistral-small-latest multimodalvisionmulti-input reasoning | Mistral AI | 38.2 overall | 31.3 | 23.2 | 0.0 | 0.0 | 81.7 | |
| 113 | Grok Code Fast 1 grok-code-fast-1 codeprogrammingtool use | xAI | 38.2 overall | 0.0 | 27.2 | 0.0 | 36.1 | 60.5 | $0.2 in / $1.5 out |
| 114 | o4-mini o4-mini multimodalvisionmulti-input reasoning | OpenAI | 37.7 overall | 46.2 | 0.0 | 36.1 | 28.7 | 0.0 | N/A |
| 115 | QvQ-72B-Preview qvq-72b-preview multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 37.6 overall | 37.6 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 116 | Gemini 3.5 Flash-Lite gemini-3.5-flash-lite multimodalvisionmulti-input reasoning | Google | 37.6 overall | 18.9 | 84.8 | 0.0 | 17.2 | 59.1 | |
| 117 | MiMo-V2-Flash mimo-v2-flash codeprogrammingtool use | Xiaomi | 37.4 overall | 50.6 | 0.0 | 23.5 | 35.8 | 0.0 | N/A |
| 118 | DeepSeek-V3.2 deepseek-v3.2 codeprogrammingtool use | DeepSeek | 37.3 overall | 55.0 | 0.0 | 12.4 | 42.0 | 0.0 | N/A |
| 119 | GLM-5 glm-5 codeprogrammingtool use | Zhipu AI | 37.3 overall | 0.0 | 6.9 | 35.1 | 60.4 | 40.0 | $1 in / $3.2 out |
| 120 | DeepSeek R1 Zero deepseek-r1-zero textinference | DeepSeek | 36.8 overall | 36.8 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
Qwen3-Next-80B-A3B-Thinking
Alibaba Cloud / Qwen Team
42.0
N/A
Qwen3 VL 235B A22B Instruct
Alibaba Cloud / Qwen Team
42.0
N/A
Qwen3-Coder 480B A35B Instruct
Alibaba Cloud / Qwen Team
41.8
N/A
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| N/A |
| N/A |
| N/A |
| N/A |
| $0.15 in / $0.6 out |
| $0.3 in / $2.5 out |
Qwen3.5-35B-A3B
Alibaba Cloud / Qwen Team
41.6
N/A
GLM-4.6
Zhipu AI
41.2
N/A
Kimi K2 0905
Moonshot AI
41.0
N/A
Qwen3-235B-A22B-Instruct-2507
Alibaba Cloud / Qwen Team
40.7
N/A
LongCat-Flash-Lite
Meituan
40.1
$0.1 in / $0.4 out
Gemini 2.5 Pro Preview 06-05
39.4
N/A
DeepSeek-V3.2-Exp
DeepSeek
39.1
N/A
LongCat-Flash-Thinking-2601
Meituan
38.3
N/A
Mistral Small 4
Mistral AI
38.2
$0.15 in / $0.6 out
Grok Code Fast 1
xAI
38.2
$0.2 in / $1.5 out
o4-mini
OpenAI
37.7
N/A
QvQ-72B-Preview
Alibaba Cloud / Qwen Team
37.6
N/A
Gemini 3.5 Flash-Lite
37.6
$0.3 in / $2.5 out
MiMo-V2-Flash
Xiaomi
37.4
N/A
DeepSeek-V3.2
DeepSeek
37.3
N/A
GLM-5
Zhipu AI
37.3
$1 in / $3.2 out
DeepSeek R1 Zero
DeepSeek
36.8
N/A