Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

28.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
201

Qwen2.5 VL 32B Instruct

qwen2.5-vl-32b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

19.4

Benchmarks

19.40.01.50.00.0N/A
202

GPT-4.1 mini

gpt-4.1-mini-2025-04-14

multimodalvisionmulti-input reasoning
OpenAI

19.2

Benchmarks

19.284.68.92.269.5
203

Nova Pro

nova-pro

multimodalvisionmulti-input reasoning
AAmazon

19.0

Benchmarks

19.00.00.00.00.0N/A
204

Gemini 3.5 Flash-Lite

gemini-3.5-flash-lite

multimodalvisionmulti-input reasoning
Google

18.9

Benchmarks

18.984.80.017.259.1
205

Mistral Small 3.2 24B Instruct

mistral-small-3.2-24b-instruct-2506

multimodalvisionmulti-input reasoning
Mistral AI

18.4

Benchmarks

18.40.00.00.00.0
206

Llama 3.1 405B Instruct

llama-3.1-405b-instruct

textinference
MMeta

18.3

Benchmarks

18.30.00.00.00.0N/A
207

Qwen3 VL 4B Instruct

qwen3-vl-4b-instruct

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

18.2

Benchmarks

18.232.917.70.085.4
208

Llama 3.3 70B Instruct

llama-3.3-70b-instruct

textinference
MMeta

18.0

Benchmarks

18.00.00.00.00.0N/A
209

Gemma 4 E4B

gemma-4-e4b-it

multimodalvisionmulti-input reasoning
Google

17.8

Benchmarks

17.80.00.00.00.0N/A
210

Claude 3 Opus

claude-3-opus-20240229

multimodalvisionmulti-input reasoning
Anthropic

17.7

Benchmarks

17.70.00.00.00.0
211

Qwen2.5 32B Instruct

qwen-2.5-32b-instruct

textinference
AAlibaba Cloud / Qwen Team

17.1

Benchmarks

17.10.00.00.00.0N/A
212

DeepSeek R1 Distill Qwen 7B

deepseek-r1-distill-qwen-7b

textinference
DeepSeek

16.8

Benchmarks

16.80.00.00.00.0N/A
213

DeepSeek R1 Distill Llama 8B

deepseek-r1-distill-llama-8b

textinference
DeepSeek

16.3

Benchmarks

16.30.00.00.00.0N/A
214

Qwen2.5 72B Instruct

qwen-2.5-72b-instruct

textinference
AAlibaba Cloud / Qwen Team

16.3

Benchmarks

16.30.00.00.00.0N/A
215

GPT-4 Turbo

gpt-4-turbo-2024-04-09

textinference
OpenAI

15.5

Benchmarks

15.550.20.00.015.4$10 in / $30 out
216

Mistral Small 3.1 24B Instruct

mistral-small-3.1-24b-instruct-2503

multimodalvisionmulti-input reasoning
Mistral AI

15.1

Benchmarks

15.10.00.00.00.0
217

Llama 3.1 Nemotron Nano 8B V1

llama-3.1-nemotron-nano-8b-v1

textinference
NNVIDIA

15.0

Benchmarks

15.00.00.00.00.0N/A
218

Llama 3.2 90B Instruct

llama-3.2-90b-instruct

multimodalvisionmulti-input reasoning
MMeta

14.9

Benchmarks

14.90.00.00.00.0N/A
219

Phi 4

phi-4

textinference
MMicrosoft

14.4

Benchmarks

14.40.00.00.00.0N/A
220

GPT-4o mini

gpt-4o-mini-2024-07-18

multimodalvisionmulti-input reasoning
OpenAI

14.2

Benchmarks

14.20.00.00.00.0
201
A

Qwen2.5 VL 32B Instruct

Alibaba Cloud / Qwen Team

19.4

N/A

202

GPT-4.1 mini

OpenAI

19.2

$0.4 in / $1.6 out

203
A

Nova Pro

Amazon

19.0

N/A

204

Page 11 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$0.4 in / $1.6 out
$0.3 in / $2.5 out
N/A
$0.1 in / $0.6 out
N/A
N/A
N/A

Gemini 3.5 Flash-Lite

Google

18.9

$0.3 in / $2.5 out

205

Mistral Small 3.2 24B Instruct

Mistral AI

18.4

N/A

206
M

Llama 3.1 405B Instruct

Meta

18.3

N/A

207
A

Qwen3 VL 4B Instruct

Alibaba Cloud / Qwen Team

18.2

$0.1 in / $0.6 out

208
M

Llama 3.3 70B Instruct

Meta

18.0

N/A

209

Gemma 4 E4B

Google

17.8

N/A

210

Claude 3 Opus

Anthropic

17.7

N/A

211
A

Qwen2.5 32B Instruct

Alibaba Cloud / Qwen Team

17.1

N/A

212

DeepSeek R1 Distill Qwen 7B

DeepSeek

16.8

N/A

213

DeepSeek R1 Distill Llama 8B

DeepSeek

16.3

N/A

214
A

Qwen2.5 72B Instruct

Alibaba Cloud / Qwen Team

16.3

N/A

215

GPT-4 Turbo

OpenAI

15.5

$10 in / $30 out

216

Mistral Small 3.1 24B Instruct

Mistral AI

15.1

N/A

217
N

Llama 3.1 Nemotron Nano 8B V1

NVIDIA

15.0

N/A

218
M

Llama 3.2 90B Instruct

Meta

14.9

N/A

219
M

Phi 4

Microsoft

14.4

N/A

220

GPT-4o mini

OpenAI

14.2

N/A