Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

28.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
221

Qwen2.5 14B Instruct

qwen-2.5-14b-instruct

textinference
AAlibaba Cloud / Qwen Team

13.4

Benchmarks

13.40.00.00.00.0N/A
222

Qwen3.5-2B

qwen3.5-2b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

13.2

Benchmarks

13.20.00.00.00.0N/A
223

Mistral Small 3 24B Instruct

mistral-small-24b-instruct-2501

textinference
Mistral AI

13.0

Benchmarks

13.00.00.00.00.0N/A
224

Mistral Small 3.1 24B Base

mistral-small-3.1-24b-base-2503

multimodalvisionmulti-input reasoning
Mistral AI

12.9

Benchmarks

12.90.00.00.00.0
225

Nova Lite

nova-lite

multimodalvisionmulti-input reasoning
AAmazon

12.9

Benchmarks

12.90.00.00.00.0N/A
226

GPT-4.1 nano

gpt-4.1-nano-2025-04-14

multimodalvisionmulti-input reasoning
OpenAI

11.6

Benchmarks

11.687.80.00.094.9
227

Qwen2 72B Instruct

qwen2-72b-instruct

textinference
AAlibaba Cloud / Qwen Team

11.0

Benchmarks

11.00.00.00.00.0N/A
228

Llama 3.1 70B Instruct

llama-3.1-70b-instruct

textinference
MMeta

10.3

Benchmarks

10.30.00.00.00.0N/A
229

Claude 3.5 Haiku

claude-3-5-haiku-20241022

codeprogrammingtool use
Anthropic

9.9

Benchmarks

9.90.03.06.60.0
230

Gemini 1.5 Flash 8B

gemini-1.5-flash-8b

multimodalvisionmulti-input reasoning
Google

9.9

Benchmarks

9.90.00.00.00.0
231

Grok-1.5V

grok-1.5v

multimodalvisionmulti-input reasoning
xAI

9.7

Benchmarks

9.70.00.00.00.0N/A
232

Claude 3 Sonnet

claude-3-sonnet-20240229

multimodalvisionmulti-input reasoning
Anthropic

9.2

Benchmarks

9.20.00.00.00.0
233

Qwen2.5 VL 7B Instruct

qwen2.5-vl-7b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

9.0

Benchmarks

9.00.00.00.00.0N/A
234

Mistral Large 3

mistral-large-3-2509

multimodalvisionmulti-input reasoning
Mistral AI

8.8

Benchmarks

8.80.00.00.00.0
235

Gemma 3 12B

gemma-3-12b-it

multimodalvisionmulti-input reasoning
Google

8.6

Benchmarks

8.60.00.00.00.0N/A
236

Gemma 4 E2B

gemma-4-e2b-it

multimodalvisionmulti-input reasoning
Google

8.5

Benchmarks

8.50.00.00.00.0N/A
237

Nova Micro

nova-micro

textinference
AAmazon

8.4

Benchmarks

8.40.00.00.00.0N/A
238

Grok-1.5

grok-1.5

multimodalvisionmulti-input reasoning
xAI

8.2

Benchmarks

8.20.00.00.00.0N/A
239

Phi-4-multimodal-instruct

phi-4-multimodal-instruct

multimodalvisionmulti-input reasoning
MMicrosoft

8.0

Benchmarks

8.00.00.00.00.0N/A
240

Pixtral-12B

pixtral-12b-2409

multimodalvisionmulti-input reasoning
Mistral AI

8.0

Benchmarks

8.00.00.00.00.0
221
A

Qwen2.5 14B Instruct

Alibaba Cloud / Qwen Team

13.4

N/A

222
A

Qwen3.5-2B

Alibaba Cloud / Qwen Team

13.2

N/A

223

Mistral Small 3 24B Instruct

Mistral AI

13.0

N/A

224

Page 12 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
$0.1 in / $0.4 out
N/A
N/A
N/A
N/A
N/A

Mistral Small 3.1 24B Base

Mistral AI

12.9

N/A

225
A

Nova Lite

Amazon

12.9

N/A

226

GPT-4.1 nano

OpenAI

11.6

$0.1 in / $0.4 out

227
A

Qwen2 72B Instruct

Alibaba Cloud / Qwen Team

11.0

N/A

228
M

Llama 3.1 70B Instruct

Meta

10.3

N/A

229

Claude 3.5 Haiku

Anthropic

9.9

N/A

230

Gemini 1.5 Flash 8B

Google

9.9

N/A

231

Grok-1.5V

xAI

9.7

N/A

232

Claude 3 Sonnet

Anthropic

9.2

N/A

233
A

Qwen2.5 VL 7B Instruct

Alibaba Cloud / Qwen Team

9.0

N/A

234

Mistral Large 3

Mistral AI

8.8

N/A

235

Gemma 3 12B

Google

8.6

N/A

236

Gemma 4 E2B

Google

8.5

N/A

237
A

Nova Micro

Amazon

8.4

N/A

238

Grok-1.5

xAI

8.2

N/A

239
M

Phi-4-multimodal-instruct

Microsoft

8.0

N/A

240

Pixtral-12B

Mistral AI

8.0

N/A