Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

12.2

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
81

MiMo-V2-Flash

mimo-v2-flash

codeprogrammingtool use
Xiaomi

23.5

Agentic

50.60.023.535.80.0N/A
82

Qwen3 VL 30B A3B Instruct

qwen3-vl-30b-a3b-instruct

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

21.4

Agentic

25.60.021.40.00.0
83

Qwen3 VL 8B Thinking

qwen3-vl-8b-thinking

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

21.1

Agentic

32.80.021.10.00.0N/A
84

MiniMax M1 80K

minimax-m1-80k

codeprogrammingtool use
MiniMax

20.9

Agentic

22.80.020.916.20.0N/A
85

Qwen3 VL 30B A3B Thinking

qwen3-vl-30b-a3b-thinking

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

19.2

Agentic

32.70.019.20.00.0
86

o3

o3-2025-04-16

multimodalvisionmulti-input reasoning
OpenAI

17.9

Agentic

42.90.017.927.70.0N/A
87

Qwen3-Next-80B-A3B-Instruct

qwen3-next-80b-a3b-instruct

textinference
AAlibaba Cloud / Qwen Team

17.9

Agentic

27.40.017.90.00.0N/A
88

Qwen3 VL 4B Instruct

qwen3-vl-4b-instruct

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

17.7

Agentic

18.232.917.70.085.4
89

Qwen3 VL 4B Thinking

qwen3-vl-4b-thinking

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

17.0

Agentic

20.232.917.00.079.3
90

Sarvam-105B

sarvam-105b

codeprogrammingtool use
SSarvam AI

16.7

Agentic

41.10.016.710.30.0N/A
91

Mistral Medium 3.5

mistral-medium-3-5

multimodalvisionmulti-input reasoning
Mistral AI

15.4

Agentic

34.623.215.459.034.8
92

GPT-5.4 Mini

gpt-5.4-mini

texttext-to-textlanguage
OpenAI

15.0

Agentic

51.144.515.020.042.7
93

GPT-4o

gpt-4o-2024-08-06

multimodalvisionmulti-input reasoning
OpenAI

14.9

Agentic

29.139.614.93.731.2
94

DeepSeek-V3.1

deepseek-v3.1

codeprogrammingtool use
DeepSeek

13.6

Agentic

36.70.013.625.40.0N/A
95

Kimi K2 Instruct

kimi-k2-instruct

codeprogrammingtool use
Moonshot AI

13.5

Agentic

23.20.013.514.00.0N/A
96

Nova 2 Lite

nova-2-lite

multimodalvisionmulti-input reasoning
AAmazon

13.0

Agentic

41.162.813.026.059.1$0.3 in / $2.5 out
97

Grok 4 Fast

grok-4-fast

multimodalvisionmulti-input reasoning
xAI

12.8

Agentic

55.70.012.80.00.0N/A
98

DeepSeek-V3.2 (Thinking)

deepseek-reasoner

codeprogrammingtool use
DeepSeek

12.4

Agentic

49.80.012.442.00.0N/A
99

DeepSeek-V3.2

deepseek-v3.2

codeprogrammingtool use
DeepSeek

12.4

Agentic

55.00.012.442.00.0N/A
100

o3-mini

o3-mini

codeprogrammingtool use
OpenAI

11.9

Agentic

25.30.011.911.60.0N/A
81

MiMo-V2-Flash

Xiaomi

23.5

N/A

82
A

Qwen3 VL 30B A3B Instruct

Alibaba Cloud / Qwen Team

21.4

N/A

83
A

Qwen3 VL 8B Thinking

Alibaba Cloud / Qwen Team

21.1

N/A

84

Page 5 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
N/A
$0.1 in / $0.6 out
$0.1 in / $1 out
$1.5 in / $7.5 out
$0.75 in / $4.5 out
$2.5 in / $10 out

MiniMax M1 80K

MiniMax

20.9

N/A

85
A

Qwen3 VL 30B A3B Thinking

Alibaba Cloud / Qwen Team

19.2

N/A

86

o3

OpenAI

17.9

N/A

87
A

Qwen3-Next-80B-A3B-Instruct

Alibaba Cloud / Qwen Team

17.9

N/A

88
A

Qwen3 VL 4B Instruct

Alibaba Cloud / Qwen Team

17.7

$0.1 in / $0.6 out

89
A

Qwen3 VL 4B Thinking

Alibaba Cloud / Qwen Team

17.0

$0.1 in / $1 out

90
S

Sarvam-105B

Sarvam AI

16.7

N/A

91

Mistral Medium 3.5

Mistral AI

15.4

$1.5 in / $7.5 out

92

GPT-5.4 Mini

OpenAI

15.0

$0.75 in / $4.5 out

93

GPT-4o

OpenAI

14.9

$2.5 in / $10 out

94

DeepSeek-V3.1

DeepSeek

13.6

N/A

95

Kimi K2 Instruct

Moonshot AI

13.5

N/A

96
A

Nova 2 Lite

Amazon

13.0

$0.3 in / $2.5 out

97

Grok 4 Fast

xAI

12.8

N/A

98

DeepSeek-V3.2 (Thinking)

DeepSeek

12.4

N/A

99

DeepSeek-V3.2

DeepSeek

12.4

N/A

100

o3-mini

OpenAI

11.9

N/A