Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

29.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
141

DeepSeek-V3.2-Speciale

deepseek-v3.2-speciale

codeprogrammingtool use
DeepSeek

33.3

overall

50.70.05.042.00.0N/A
142

Llama 3.1 Nemotron Ultra 253B v1

llama-3.1-nemotron-ultra-253b-v1

textinference
NNVIDIA

33.0

overall

33.00.00.00.00.0N/A
143

Gemma 4 12B

gemma-4-12b-it

multimodalvisionmulti-input reasoning
Google

32.6

overall

32.60.00.00.00.0N/A
144

Llama 4 Maverick

llama-4-maverick

multimodalvisionmulti-input reasoning
MMeta

32.6

overall

32.60.00.00.00.0N/A
145

MiniMax M2.7

minimax-m2.7

codeprogrammingtool use
MiniMax

32.1

overall

0.019.526.329.073.2$0.3 in / $1.2 out
146

o1

o1-2024-12-17

multimodalvisionmulti-input reasoning
OpenAI

32.0

overall

41.90.044.75.60.0N/A
147

Gemini 2.0 Flash

gemini-2.0-flash

multimodalvisionmulti-input reasoning
Google

31.7

overall

31.70.00.00.00.0
148

GPT-5.3 Codex

gpt-5.3-codex

texttext-to-textcoding
OpenAI

30.7

overall

0.031.10.034.522.0
149

GPT OSS 120B

gpt-oss-120b

textinference
OpenAI

30.5

overall

33.70.026.80.00.0N/A
150

DeepSeek-V3 0324

deepseek-v3-0324

textinference
DeepSeek

30.4

overall

30.40.00.00.00.0N/A
151

o3

o3-2025-04-16

multimodalvisionmulti-input reasoning
OpenAI

30.3

overall

42.90.017.927.70.0N/A
152

Qwen3.6-35B-A3B

qwen3.6-35b-a3b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

30.0

overall

51.10.09.825.20.0N/A
153

GPT-5.4 nano

gpt-5.4-nano

multimodalvisionmulti-input reasoning
OpenAI

30.0

overall

41.844.56.98.276.8$0.2 in / $1.25 out
154

Gemini 2.5 Flash

gemini-2.5-flash

multimodalvisionmulti-input reasoning
Google

30.0

overall

38.20.00.019.50.0
155

Qwen3.5-4B

qwen3.5-4b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

29.9

overall

29.90.00.00.00.0N/A
156

Qwen3 Max

qwen3-max

codeprogrammingtool use
AAlibaba Cloud / Qwen Team

29.8

overall

28.00.00.032.10.0N/A
157

Qwen3 VL 4B Instruct

qwen3-vl-4b-instruct

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

29.7

overall

18.232.917.70.085.4
158

Ministral 3 (8B Reasoning 2512)

ministral-8b-latest

multimodalvisionmulti-input reasoning
Mistral AI

29.4

overall

29.40.00.00.00.0
159

Phi 4 Reasoning Plus

phi-4-reasoning-plus

textinference
MMicrosoft

29.4

overall

29.40.00.00.00.0N/A
160

Qwen3 VL 4B Thinking

qwen3-vl-4b-thinking

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

29.4

overall

20.232.917.00.079.3
141

DeepSeek-V3.2-Speciale

DeepSeek

33.3

N/A

142
N

Llama 3.1 Nemotron Ultra 253B v1

NVIDIA

33.0

N/A

143

Gemma 4 12B

Google

32.6

N/A

144

Page 8 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
$1.75 in / $14 out
N/A
$0.1 in / $0.6 out
N/A
$0.1 in / $1 out
M

Llama 4 Maverick

Meta

32.6

N/A

145

MiniMax M2.7

MiniMax

32.1

$0.3 in / $1.2 out

146

o1

OpenAI

32.0

N/A

147

Gemini 2.0 Flash

Google

31.7

N/A

148

GPT-5.3 Codex

OpenAI

30.7

$1.75 in / $14 out

149

GPT OSS 120B

OpenAI

30.5

N/A

150

DeepSeek-V3 0324

DeepSeek

30.4

N/A

151

o3

OpenAI

30.3

N/A

152
A

Qwen3.6-35B-A3B

Alibaba Cloud / Qwen Team

30.0

N/A

153

GPT-5.4 nano

OpenAI

30.0

$0.2 in / $1.25 out

154

Gemini 2.5 Flash

Google

30.0

N/A

155
A

Qwen3.5-4B

Alibaba Cloud / Qwen Team

29.9

N/A

156
A

Qwen3 Max

Alibaba Cloud / Qwen Team

29.8

N/A

157
A

Qwen3 VL 4B Instruct

Alibaba Cloud / Qwen Team

29.7

$0.1 in / $0.6 out

158

Ministral 3 (8B Reasoning 2512)

Mistral AI

29.4

N/A

159
M

Phi 4 Reasoning Plus

Microsoft

29.4

N/A

160
A

Qwen3 VL 4B Thinking

Alibaba Cloud / Qwen Team

29.4

$0.1 in / $1 out