Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

29.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
201

Magistral Small 2506

magistral-small-2506

textinference
Mistral AI

22.7

overall

22.70.00.00.00.0N/A
202

GLM-4.5-Air

glm-4.5-air

codeprogrammingtool use
ZZhipu AI

22.3

overall

25.70.024.216.00.0N/A
203

GPT-4o

gpt-4o-2024-08-06

multimodalvisionmulti-input reasoning
OpenAI

22.0

overall

29.139.614.93.731.2
204

GLM-4.7-Flash

glm-4.7-flash

codeprogrammingtool use
ZZhipu AI

21.9

overall

36.40.09.017.70.0N/A
205

Mistral Large 3 (675B Base)

mistral-large-3-675b-base-2512

multimodalvisionmulti-input reasoning
Mistral AI

21.9

overall

21.90.00.00.00.0
206

Mistral Large 3 (675B Instruct 2512 Eagle)

mistral-large-3-675B-instruct-2512-eagle

multimodalvisionmulti-input reasoning
Mistral AI

21.9

overall

21.90.00.00.00.0
207

Mistral Large 3 (675B Instruct 2512 NVFP4)

mistral-large-3-675b-instruct-2512-nvfp4

multimodalvisionmulti-input reasoning
Mistral AI

21.9

overall

21.90.00.00.00.0
208

Gemini 1.5 Flash

gemini-1.5-flash

multimodalvisionmulti-input reasoning
Google

21.8

overall

21.80.00.00.00.0
209

MiniMax M1 40K

minimax-m1-40k

codeprogrammingtool use
MiniMax

21.4

overall

21.30.026.815.50.0N/A
210

Llama-3.3 Nemotron Super 49B v1

llama-3.3-nemotron-super-49b-v1

textinference
NNVIDIA

21.3

overall

21.30.00.00.00.0N/A
211

Phi 4 Reasoning

phi-4-reasoning

textinference
MMicrosoft

21.3

overall

21.30.00.00.00.0N/A
212

GPT-3.5 Turbo

gpt-3.5-turbo-0125

multimodalvisionmulti-input reasoning
OpenAI

21.3

overall

2.330.10.00.060.7
213

Magistral Medium

magistral-medium

multimodalvisionmulti-input reasoning
Mistral AI

20.6

overall

20.60.00.00.00.0
214

Devstral Medium

devstral-medium-2507

codeprogrammingtool use
Mistral AI

20.6

overall

0.00.00.020.60.0N/A
215

Sarvam-30B

sarvam-30b

codeprogrammingtool use
SSarvam AI

20.4

overall

44.90.06.44.40.0N/A
216

Min istral 3 (3B Reasoning 2512)

ministral-3b-latest

multimodalvisionmulti-input reasoning
Mistral AI

20.4

overall

20.40.00.00.00.0
217

MiniMax M1 80K

minimax-m1-80k

codeprogrammingtool use
MiniMax

20.2

overall

22.80.020.916.20.0N/A
218

GPT-5 nano

gpt-5-nano-2025-08-07

multimodalvisionmulti-input reasoning
OpenAI

20.0

overall

25.20.00.013.30.0
219

Phi 4 Mini Reasoning

phi-4-mini-reasoning

textinference
MMicrosoft

19.9

overall

19.90.00.00.00.0N/A
220

DeepSeek-R1-0528

deepseek-r1-0528

codeprogrammingtool use
DeepSeek

19.8

overall

47.90.00.05.70.0N/A
201

Magistral Small 2506

Mistral AI

22.7

N/A

202
Z

GLM-4.5-Air

Zhipu AI

22.3

N/A

203

GPT-4o

OpenAI

22.0

$2.5 in / $10 out

204

Page 11 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$2.5 in / $10 out
N/A
N/A
N/A
N/A
$0.5 in / $1.5 out
N/A
N/A
N/A
Z

GLM-4.7-Flash

Zhipu AI

21.9

N/A

205

Mistral Large 3 (675B Base)

Mistral AI

21.9

N/A

206

Mistral Large 3 (675B Instruct 2512 Eagle)

Mistral AI

21.9

N/A

207

Mistral Large 3 (675B Instruct 2512 NVFP4)

Mistral AI

21.9

N/A

208

Gemini 1.5 Flash

Google

21.8

N/A

209

MiniMax M1 40K

MiniMax

21.4

N/A

210
N

Llama-3.3 Nemotron Super 49B v1

NVIDIA

21.3

N/A

211
M

Phi 4 Reasoning

Microsoft

21.3

N/A

212

GPT-3.5 Turbo

OpenAI

21.3

$0.5 in / $1.5 out

213

Magistral Medium

Mistral AI

20.6

N/A

214

Devstral Medium

Mistral AI

20.6

N/A

215
S

Sarvam-30B

Sarvam AI

20.4

N/A

216

Min istral 3 (3B Reasoning 2512)

Mistral AI

20.4

N/A

217

MiniMax M1 80K

MiniMax

20.2

N/A

218

GPT-5 nano

OpenAI

20.0

N/A

219
M

Phi 4 Mini Reasoning

Microsoft

19.9

N/A

220

DeepSeek-R1-0528

DeepSeek

19.8

N/A