Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

28.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
181

LongCat-Flash-Lite

longcat-flash-lite

codeprogrammingtool use
Meituan

22.8

Benchmarks

22.872.830.123.995.6$0.1 in / $0.4 out
182

MiniMax M1 80K

minimax-m1-80k

codeprogrammingtool use
MiniMax

22.8

Benchmarks

22.80.020.916.20.0N/A
183

DeepSeek R1 Distill Qwen 14B

deepseek-r1-distill-qwen-14b

textinference
DeepSeek

22.7

Benchmarks

22.70.00.00.00.0N/A
184

Magistral Small 2506

magistral-small-2506

textinference
Mistral AI

22.7

Benchmarks

22.70.00.00.00.0N/A
185

Qwen2.5 VL 72B Instruct

qwen2.5-vl-72b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

22.7

Benchmarks

22.70.05.20.00.0N/A
186

Mistral Large 3 (675B Base)

mistral-large-3-675b-base-2512

multimodalvisionmulti-input reasoning
Mistral AI

21.9

Benchmarks

21.90.00.00.00.0
187

Mistral Large 3 (675B Instruct 2512 Eagle)

mistral-large-3-675B-instruct-2512-eagle

multimodalvisionmulti-input reasoning
Mistral AI

21.9

Benchmarks

21.90.00.00.00.0
188

Mistral Large 3 (675B Instruct 2512 NVFP4)

mistral-large-3-675b-instruct-2512-nvfp4

multimodalvisionmulti-input reasoning
Mistral AI

21.9

Benchmarks

21.90.00.00.00.0
189

Mistral Large 3 (675B Instruct 2512)

mistral-large-latest

multimodalvisionmulti-input reasoning
Mistral AI

21.9

Benchmarks

21.922.00.00.055.6
190

Gemini 1.5 Flash

gemini-1.5-flash

multimodalvisionmulti-input reasoning
Google

21.8

Benchmarks

21.80.00.00.00.0
191

Llama-3.3 Nemotron Super 49B v1

llama-3.3-nemotron-super-49b-v1

textinference
NNVIDIA

21.3

Benchmarks

21.30.00.00.00.0N/A
192

MiniMax M1 40K

minimax-m1-40k

codeprogrammingtool use
MiniMax

21.3

Benchmarks

21.30.026.815.50.0N/A
193

Phi 4 Reasoning

phi-4-reasoning

textinference
MMicrosoft

21.3

Benchmarks

21.30.00.00.00.0N/A
194

Magistral Medium

magistral-medium

multimodalvisionmulti-input reasoning
Mistral AI

20.6

Benchmarks

20.60.00.00.00.0
195

GPT-4o

gpt-4o-2024-05-13

multimodalvisionmulti-input reasoning
OpenAI

20.5

Benchmarks

20.538.00.00.030.7
196

Min istral 3 (3B Reasoning 2512)

ministral-3b-latest

multimodalvisionmulti-input reasoning
Mistral AI

20.4

Benchmarks

20.40.00.00.00.0
197

Gemini 2.5 Flash-Lite

gemini-2.5-flash-lite

multimodalvisionmulti-input reasoning
Google

20.3

Benchmarks

20.30.00.02.90.0
198

Qwen3 VL 4B Thinking

qwen3-vl-4b-thinking

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

20.2

Benchmarks

20.232.917.00.079.3
199

Qwen3 32B

qwen3-32b

textinference
AAlibaba Cloud / Qwen Team

20.1

Benchmarks

20.12.20.00.078.0$0.1 in / $0.3 out
200

Phi 4 Mini Reasoning

phi-4-mini-reasoning

textinference
MMicrosoft

19.9

Benchmarks

19.90.00.00.00.0N/A
181

LongCat-Flash-Lite

Meituan

22.8

$0.1 in / $0.4 out

182

MiniMax M1 80K

MiniMax

22.8

N/A

183

DeepSeek R1 Distill Qwen 14B

DeepSeek

22.7

N/A

184

Page 10 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
N/A
N/A
$0.5 in / $1.5 out
N/A
N/A
$2.5 in / $10 out
N/A
N/A
$0.1 in / $1 out

Magistral Small 2506

Mistral AI

22.7

N/A

185
A

Qwen2.5 VL 72B Instruct

Alibaba Cloud / Qwen Team

22.7

N/A

186

Mistral Large 3 (675B Base)

Mistral AI

21.9

N/A

187

Mistral Large 3 (675B Instruct 2512 Eagle)

Mistral AI

21.9

N/A

188

Mistral Large 3 (675B Instruct 2512 NVFP4)

Mistral AI

21.9

N/A

189

Mistral Large 3 (675B Instruct 2512)

Mistral AI

21.9

$0.5 in / $1.5 out

190

Gemini 1.5 Flash

Google

21.8

N/A

191
N

Llama-3.3 Nemotron Super 49B v1

NVIDIA

21.3

N/A

192

MiniMax M1 40K

MiniMax

21.3

N/A

193
M

Phi 4 Reasoning

Microsoft

21.3

N/A

194

Magistral Medium

Mistral AI

20.6

N/A

195

GPT-4o

OpenAI

20.5

$2.5 in / $10 out

196

Min istral 3 (3B Reasoning 2512)

Mistral AI

20.4

N/A

197

Gemini 2.5 Flash-Lite

Google

20.3

N/A

198
A

Qwen3 VL 4B Thinking

Alibaba Cloud / Qwen Team

20.2

$0.1 in / $1 out

199
A

Qwen3 32B

Alibaba Cloud / Qwen Team

20.1

$0.1 in / $0.3 out

200
M

Phi 4 Mini Reasoning

Microsoft

19.9

N/A