Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.
334
Tracked models
29
Providers
286
Benchmarked
28.3
Avg. index
334 models
| Rank | Model | Provider | Score | Benchmarks | Inference | Agentic | Programming | Value | Price |
|---|---|---|---|---|---|---|---|---|---|
| 61 | GPT-5 Medium gpt-5-medium-2025-08-07 multimodalvisionmulti-input reasoning | OpenAI | 54.6 Benchmarks | 54.6 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 62 | Claude Opus 4.5 claude-opus-4-5-20251101 multimodalvisionmulti-input reasoning | Anthropic | 54.5 Benchmarks | 54.5 | 0.0 | 35.2 | 72.2 | 0.0 | |
| 63 | Qwen3.5-397B-A17B qwen3.5-397b-a17b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 54.5 Benchmarks | 54.5 | 0.0 | 24.7 | 55.3 | 0.0 | N/A |
| 64 | Qwen3.5-35B-A3B qwen3.5-35b-a3b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 53.4 Benchmarks | 53.4 | 0.0 | 38.3 | 30.2 | 0.0 | N/A |
| 65 | ChatGPT-4o Latest chatgpt-4o-latest multimodalvisionmulti-input reasoning | OpenAI | 53.0 Benchmarks | 53.0 | 0.0 | 0.0 | 0.0 | 0.0 | |
| 66 | LongCat-Flash-Thinking-2601 longcat-flash-thinking-2601 codeprogrammingtool use | Meituan | 52.9 Benchmarks | 52.9 | 0.0 | 25.6 | 33.5 | 0.0 | |
| 67 | Gemini 3.1 Flash-Lite gemini-3.1-flash-lite-preview multimodalvisionmulti-input reasoning | Google | 52.6 Benchmarks | 52.6 | 62.8 | 0.0 | 0.0 | 67.1 | |
| 68 | GPT OSS 20B High gpt-oss-20b-high textinference | OpenAI | 52.3 Benchmarks | 52.3 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 69 | Claude Sonnet 4.5 claude-sonnet-4-5-20250929 multimodalvisionmulti-input reasoning | Anthropic | 51.4 Benchmarks | 51.4 | 12.6 | 69.9 | 74.6 | 12.0 | |
| 70 | GPT-5.4 Mini gpt-5.4-mini texttext-to-textlanguage | OpenAI | 51.1 Benchmarks | 51.1 | 44.5 | 15.0 | 20.0 | 42.7 | |
| 71 | Qwen3.6-35B-A3B qwen3.6-35b-a3b multimodalvisionmulti-input reasoning | Alibaba Cloud / Qwen Team | 51.1 Benchmarks | 51.1 | 0.0 | 9.8 | 25.2 | 0.0 | N/A |
| 72 | Grok-3 Mini grok-3-mini multimodalvisionmulti-input reasoning | xAI | 51.0 Benchmarks | 51.0 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
| 73 | DeepSeek-V3.2-Speciale deepseek-v3.2-speciale codeprogrammingtool use | DeepSeek | 50.7 Benchmarks | 50.7 | 0.0 | 5.0 | 42.0 | 0.0 | |
| 74 | MiMo-V2-Flash mimo-v2-flash codeprogrammingtool use | Xiaomi | 50.6 Benchmarks | 50.6 | 0.0 | 23.5 | 35.8 | 0.0 | N/A |
| 75 | DeepSeek-V3.2-Exp deepseek-v3.2-exp codeprogrammingtool use | DeepSeek | 50.2 Benchmarks | 50.2 | 0.0 | 27.2 | 38.0 | 0.0 | N/A |
| 76 | Gemini 2.5 Pro Preview 06-05 gemini-2.5-pro-preview-06-05 multimodalvisionmulti-input reasoning | Google | 50.2 Benchmarks | 50.2 | 0.0 | 0.0 | 25.8 | 0.0 | |
| 77 | DeepSeek-V3.2 (Thinking) deepseek-reasoner codeprogrammingtool use | DeepSeek | 49.8 Benchmarks | 49.8 | 0.0 | 12.4 | 42.0 | 0.0 | |
| 78 | MiniMax M3 minimax-m3 multimodalvisionmulti-input reasoning | MiniMax | 49.6 Benchmarks | 49.6 | 62.8 | 37.5 | 68.3 | 73.2 | $0.3 in / $1.2 out |
| 79 | GPT-5.5 Instant gpt-5.5-instant multimodalvisionmulti-input reasoning | OpenAI | 49.5 Benchmarks | 49.5 | 62.8 | 0.0 | 0.0 | 17.3 | |
| 80 | Grok-4 grok-4 multimodalvisionmulti-input reasoning | xAI | 49.1 Benchmarks | 49.1 | 0.0 | 0.0 | 0.0 | 0.0 | N/A |
GPT-5 Medium
OpenAI
54.6
N/A
Claude Opus 4.5
Anthropic
54.5
N/A
Qwen3.5-397B-A17B
Alibaba Cloud / Qwen Team
54.5
N/A
Want benchmark charts, model comparison, and pricing analytics?
Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.
Open full leaderboardRankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.
| N/A |
| N/A |
| N/A |
| $0.25 in / $1.5 out |
| $3 in / $15 out |
| $0.75 in / $4.5 out |
| N/A |
| N/A |
| N/A |
| $5 in / $30 out |
Qwen3.5-35B-A3B
Alibaba Cloud / Qwen Team
53.4
N/A
ChatGPT-4o Latest
OpenAI
53.0
N/A
LongCat-Flash-Thinking-2601
Meituan
52.9
N/A
Gemini 3.1 Flash-Lite
52.6
$0.25 in / $1.5 out
GPT OSS 20B High
OpenAI
52.3
N/A
Claude Sonnet 4.5
Anthropic
51.4
$3 in / $15 out
Qwen3.6-35B-A3B
Alibaba Cloud / Qwen Team
51.1
N/A
Grok-3 Mini
xAI
51.0
N/A
DeepSeek-V3.2-Speciale
DeepSeek
50.7
N/A
MiMo-V2-Flash
Xiaomi
50.6
N/A
DeepSeek-V3.2-Exp
DeepSeek
50.2
N/A
Gemini 2.5 Pro Preview 06-05
50.2
N/A
DeepSeek-V3.2 (Thinking)
DeepSeek
49.8
N/A
MiniMax M3
MiniMax
49.6
$0.3 in / $1.2 out
GPT-5.5 Instant
OpenAI
49.5
$5 in / $30 out
Grok-4
xAI
49.1
N/A