Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
15.1
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 61 | DeepSeek-V3.2-Speciale deepseek-v3.2-speciale codeprogrammingtool use | DeepSeek | 42.0 Programming | 50.7 | 0.0 | 5.0 | 42.0 | 0.0 | N/A |
| 62 | Kimi K2.5 kimi-k2.5 multimodalvisionmulti-input reasoning | Moonshot AI | 41.8 Programming | 63.7 | 0.0 | 41.0 | 41.8 | 0.0 | N/A |
| 63 | DeepSeek-V4-Flash-Max deepseek-v4-flash-max codeprogrammingtool use | DeepSeek | 41.4 Programming | 56.2 | 84.8 | 35.3 | 41.4 | 98.8 | |
| 64 | Nemotron 3 Ultra (550B A55B) nemotron-3-ultra-550b-a55b codeprogrammingtool use | NVIDIA | 40.6 Programming | 54.7 | 0.0 | 11.5 | 40.6 | 0.0 | N/A |
| 65 | MiniMax M2 minimax-m2 codeprogrammingtool use | MiniMax | 40.0 Programming | 29.8 | 62.8 | 40.2 | 40.0 | 73.2 | $0.3 in / $1.2 out |
| 66 | Qwen3.6-27B qwen3.6-27b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 39.4 Programming | 54.7 | 32.9 | 0.0 | 39.4 | 49.4 | $0.6 in / $3.6 out |
| 67 | Qwen3.5-27B qwen3.5-27b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 39.1 Programming | 57.8 | 32.9 | 41.3 | 39.1 | 61.0 | $0.3 in / $2.4 out |
| 68 | Qwen3.5-122B-A10B qwen3.5-122b-a10b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 38.3 Programming | 60.5 | 0.0 | 44.6 | 38.3 | 0.0 | N/A |
| 69 | DeepSeek-V3.2-Exp deepseek-v3.2-exp codeprogrammingtool use | DeepSeek | 38.0 Programming | 50.2 | 0.0 | 27.2 | 38.0 | 0.0 | N/A |
| 70 | Claude 3.7 Sonnet claude-3-7-sonnet-20250219 multimodalvisionmulti-input reasoning | Anthropic | 37.7 Programming | 42.3 | 0.0 | 49.1 | 37.7 | 0.0 | |
| 71 | LongCat-Flash-Chat longcat-flash-chat codeprogrammingtool use | Meituan | 36.6 Programming | 26.0 | 0.0 | 48.1 | 36.6 | 0.0 | N/A |
| 72 | Muse Spark muse-spark multimodalvisionmulti-input reasoning | Meta | 36.2 Programming | 67.1 | 0.0 | 64.4 | 36.2 | 0.0 | N/A |
| 73 | GLM-4.5 glm-4.5 codeprogrammingtool use | Zhipu AI | 36.1 Programming | 31.5 | 0.0 | 36.0 | 36.1 | 0.0 | N/A |
| 74 | Grok Code Fast 1 grok-code-fast-1 codeprogrammingtool use | xAI | 36.1 Programming | 0.0 | 27.2 | 0.0 | 36.1 | 60.5 | $0.2 in / $1.5 out |
| 75 | MiMo-V2-Flash mimo-v2-flash codeprogrammingtool use | Xiaomi | 35.8 Programming | 50.6 | 0.0 | 23.5 | 35.8 | 0.0 | N/A |
| 76 | GPT-5.3 Codex gpt-5.3-codex texttext-to-textcoding | OpenAI | 34.5 Programming | 0.0 | 31.1 | 0.0 | 34.5 | 22.0 | |
| 77 | LongCat-Flash-Thinking-2601 longcat-flash-thinking-2601 codeprogrammingtool use | Meituan | 33.5 Programming | 52.9 | 0.0 | 25.6 | 33.5 | 0.0 | |
| 78 | MAI-Thinking-1 mai-thinking-1 codeprogrammingtool use | Microsoft | 32.2 Programming | 60.1 | 0.0 | 0.0 | 32.2 | 0.0 | N/A |
| 79 | Qwen3-Coder 480B A35B Instruct qwen3-coder-480b-a35b-instruct codeprogrammingtool use | Alibaba Cloud / Qwen Team | 32.1 Programming | 0.0 | 0.0 | 50.7 | 32.1 | 0.0 | |
| 80 | Qwen3 Max qwen3-max codeprogrammingtool use | Alibaba Cloud / Qwen Team | 32.1 Programming | 28.0 | 0.0 | 0.0 | 32.1 | 0.0 | N/A |
DeepSeek-V3.2-Speciale
DeepSeek
42.0
N/A
Kimi K2.5
Moonshot AI
41.8
N/A
DeepSeek-V4-Flash-Max
DeepSeek
41.4
$0.1 in / $0.2 out
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| $0.1 in / $0.2 out |
| N/A |
| $1.75 in / $14 out |
| N/A |
| N/A |
Nemotron 3 Ultra (550B A55B)
NVIDIA
40.6
N/A
MiniMax M2
MiniMax
40.0
$0.3 in / $1.2 out
Qwen3.6-27B
Alibaba Cloud / Qwen Team
39.4
$0.6 in / $3.6 out
Qwen3.5-27B
Alibaba Cloud / Qwen Team
39.1
$0.3 in / $2.4 out
Qwen3.5-122B-A10B
Alibaba Cloud / Qwen Team
38.3
N/A
DeepSeek-V3.2-Exp
DeepSeek
38.0
N/A
Claude 3.7 Sonnet
Anthropic
37.7
N/A
LongCat-Flash-Chat
Meituan
36.6
N/A
Muse Spark
Meta
36.2
N/A
GLM-4.5
Zhipu AI
36.1
N/A
Grok Code Fast 1
xAI
36.1
$0.2 in / $1.5 out
MiMo-V2-Flash
Xiaomi
35.8
N/A
LongCat-Flash-Thinking-2601
Meituan
33.5
N/A
MAI-Thinking-1
Microsoft
32.2
N/A
Qwen3-Coder 480B A35B Instruct
Alibaba Cloud / Qwen Team
32.1
N/A
Qwen3 Max
Alibaba Cloud / Qwen Team
32.1
N/A