Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

28.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
241

Qwen3 VL 8B Instruct

qwen3-vl-8b-instruct

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

8.0

Benchmarks

8.00.024.00.00.0N/A
242

Gemma 3 27B

gemma-3-27b-it

multimodalvisionmulti-input reasoning
Google

7.6

Benchmarks

7.60.00.00.00.0N/A
243

Jamba 1.5 Large

jamba-1.5-large

textinference
AAI21 Labs

7.5

Benchmarks

7.50.00.00.00.0N/A
244

Phi-3.5-MoE-instruct

phi-3.5-moe-instruct

multimodalvisionmulti-input reasoning
MMicrosoft

7.5

Benchmarks

7.50.00.00.00.0N/A
245

Qwen2.5-Omni-7B

qwen2.5-omni-7b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

7.2

Benchmarks

7.20.00.00.00.0N/A
246

DeepSeek VL2

deepseek-vl2

multimodalvisionmulti-input reasoning
DeepSeek

6.8

Benchmarks

6.80.00.00.00.0N/A
247

Qwen2.5 7B Instruct

qwen-2.5-7b-instruct

textinference
AAlibaba Cloud / Qwen Team

6.8

Benchmarks

6.80.00.00.00.0N/A
248

Qwen2-VL-72B-Instruct

qwen2-vl-72b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

6.8

Benchmarks

6.80.00.00.00.0N/A
249

Gemini Diffusion

gemini-diffusion

codeprogrammingtool use
Google

6.5

Benchmarks

6.50.00.01.50.0N/A
250

GPT-4

gpt-4-0613

multimodalvisionmulti-input reasoning
OpenAI

6.2

Benchmarks

6.20.00.00.00.0N/A
251

Mistral Small 3 24B Base

mistral-small-24b-base-2501

multimodalvisionmulti-input reasoning
Mistral AI

5.9

Benchmarks

5.90.00.00.00.0
252

DeepSeek R1 Distill Qwen 1.5B

deepseek-r1-distill-qwen-1.5b

textinference
DeepSeek

5.6

Benchmarks

5.60.00.00.00.0N/A
253

Claude 3 Haiku

claude-3-haiku-20240307

multimodalvisionmulti-input reasoning
Anthropic

5.3

Benchmarks

5.30.00.00.00.0
254

Llama 3.2 3B Instruct

llama-3.2-3b-instruct

textinference
MMeta

4.8

Benchmarks

4.80.00.00.00.0N/A
255

DeepSeek VL2 Small

deepseek-vl2-small

multimodalvisionmulti-input reasoning
DeepSeek

4.6

Benchmarks

4.60.00.00.00.0
256

Gemma 3 4B

gemma-3-4b-it

multimodalvisionmulti-input reasoning
Google

4.3

Benchmarks

4.30.00.00.00.0N/A
257

Jamba 1.5 Mini

jamba-1.5-mini

textinference
AAI21 Labs

4.3

Benchmarks

4.30.00.00.00.0N/A
258

GPT-5.1 Codex Mini

gpt-5.1-codex-mini

multimodalvisionmulti-input reasoning
OpenAI

3.8

Benchmarks

3.80.00.00.00.0
259

Llama 3.2 11B Instruct

llama-3.2-11b-instruct

multimodalvisionmulti-input reasoning
MMeta

3.8

Benchmarks

3.80.00.00.00.0N/A
260

Gemini 1.0 Pro

gemini-1.0-pro

multimodalvisionmulti-input reasoning
Google

3.0

Benchmarks

3.00.00.00.00.0
241
A

Qwen3 VL 8B Instruct

Alibaba Cloud / Qwen Team

8.0

N/A

242

Gemma 3 27B

Google

7.6

N/A

243
A

Jamba 1.5 Large

AI21 Labs

7.5

N/A

244

Page 13 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
N/A
N/A
N/A
N/A
M

Phi-3.5-MoE-instruct

Microsoft

7.5

N/A

245
A

Qwen2.5-Omni-7B

Alibaba Cloud / Qwen Team

7.2

N/A

246

DeepSeek VL2

DeepSeek

6.8

N/A

247
A

Qwen2.5 7B Instruct

Alibaba Cloud / Qwen Team

6.8

N/A

248
A

Qwen2-VL-72B-Instruct

Alibaba Cloud / Qwen Team

6.8

N/A

249

Gemini Diffusion

Google

6.5

N/A

250

GPT-4

OpenAI

6.2

N/A

251

Mistral Small 3 24B Base

Mistral AI

5.9

N/A

252

DeepSeek R1 Distill Qwen 1.5B

DeepSeek

5.6

N/A

253

Claude 3 Haiku

Anthropic

5.3

N/A

254
M

Llama 3.2 3B Instruct

Meta

4.8

N/A

255

DeepSeek VL2 Small

DeepSeek

4.6

N/A

256

Gemma 3 4B

Google

4.3

N/A

257
A

Jamba 1.5 Mini

AI21 Labs

4.3

N/A

258

GPT-5.1 Codex Mini

OpenAI

3.8

N/A

259
M

Llama 3.2 11B Instruct

Meta

3.8

N/A

260

Gemini 1.0 Pro

Google

3.0

N/A