Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

29.3

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
21

Gemini 3 Pro

gemini-3-pro-preview

multimodalvisionmulti-input reasoning
Google

61.0

overall

70.30.058.052.30.0N/A
22

GPT-5.1 Codex High

gpt-5.1-codex-high

multimodalvisionmulti-input reasoning
OpenAI

61.0

overall

61.00.00.00.00.0
23

Qwen3.7 Max

qwen3.7-max

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

60.8

overall

64.762.849.674.243.9$1.25 in / $3.75 out
24

Nova 2 Sonic

nova-2-sonic

multimodalvisionmulti-input reasoning
AAmazon

60.7

overall

0.062.80.00.057.3$0.33 in / $2.75 out
25

GPT-5.5 Pro

gpt-5.5-pro

multimodalvisionmulti-input reasoning
OpenAI

60.1

overall

61.10.069.248.80.0N/A
26

DeepSeek-V4-Pro-Max

deepseek-v4-pro-max

codeprogrammingtool use
DeepSeek

59.9

overall

64.984.851.052.045.1
27

Qwen3.7-Plus

qwen3.7-plus

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

59.6

overall

62.962.848.161.169.5$0.32 in / $1.28 out
28

Gemini 3.6 Flash

gemini-3.6-flash

multimodalvisionmulti-input reasoning
Google

59.0

overall

57.884.80.052.634.8
29

Kimi K2-Thinking-0905

kimi-k2-thinking-0905

codeprogrammingtool use
Moonshot AI

58.8

overall

66.20.050.858.00.0
30

Grok 4.3

grok-4.3

textinference
xAI

58.8

overall

0.062.80.00.052.4$1.25 in / $2.5 out
31

Grok 4.5

grok-4.5

multimodalvisionmulti-input reasoning
xAI

58.7

overall

69.338.20.070.835.6$2 in / $6 out
32

Gemini 3.1 Flash-Lite

gemini-3.1-flash-lite-preview

multimodalvisionmulti-input reasoning
Google

58.3

overall

52.662.80.00.067.1
33

Gemini 3.1 Pro

gemini-3.1-pro-preview

multimodalvisionmulti-input reasoning
Google

57.9

overall

72.358.352.660.922.7
34

MiMo-V2.5-Pro

mimo-v2.5-pro

codeprogrammingtool use
Xiaomi

57.7

overall

36.284.80.056.378.0$0.435 in / $0.87 out
35

GPT-5.1 Thinking

gpt-5.1-thinking-2025-11-12

multimodalvisionmulti-input reasoning
OpenAI

57.6

overall

62.40.00.051.40.0
36

Claude Opus 4.1

claude-opus-4-1-20250805

multimodalvisionmulti-input reasoning
Anthropic

57.3

overall

46.00.067.460.80.0
37

Muse Spark

muse-spark

multimodalvisionmulti-input reasoning
MMeta

57.0

overall

67.10.064.436.20.0N/A
38

ERNIE 5.0

ernie-5.0

multimodalvisionmulti-input reasoning
BBaidu

56.8

overall

56.80.00.00.00.0N/A
39

DeepSeek-V4-Flash-Max

deepseek-v4-flash-max

codeprogrammingtool use
DeepSeek

56.7

overall

56.284.835.341.498.8
40

Claude Opus 4.7

claude-opus-4-7

multimodalvisionmulti-input reasoning
Anthropic

56.5

overall

75.128.354.277.98.0
21

Gemini 3 Pro

Google

61.0

N/A

22

GPT-5.1 Codex High

OpenAI

61.0

N/A

23
A

Qwen3.7 Max

Alibaba Cloud / Qwen Team

60.8

$1.25 in / $3.75 out

24

Page 2 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
$1.6 in / $3.2 out
$1.5 in / $7.5 out
N/A
$0.25 in / $1.5 out
$2.5 in / $15 out
N/A
N/A
$0.1 in / $0.2 out
$5 in / $25 out
A

Nova 2 Sonic

Amazon

60.7

$0.33 in / $2.75 out

25

GPT-5.5 Pro

OpenAI

60.1

N/A

26

DeepSeek-V4-Pro-Max

DeepSeek

59.9

$1.6 in / $3.2 out

27
A

Qwen3.7-Plus

Alibaba Cloud / Qwen Team

59.6

$0.32 in / $1.28 out

28

Gemini 3.6 Flash

Google

59.0

$1.5 in / $7.5 out

29

Kimi K2-Thinking-0905

Moonshot AI

58.8

N/A

30

Grok 4.3

xAI

58.8

$1.25 in / $2.5 out

31

Grok 4.5

xAI

58.7

$2 in / $6 out

32

Gemini 3.1 Flash-Lite

Google

58.3

$0.25 in / $1.5 out

33

Gemini 3.1 Pro

Google

57.9

$2.5 in / $15 out

34

MiMo-V2.5-Pro

Xiaomi

57.7

$0.435 in / $0.87 out

35

GPT-5.1 Thinking

OpenAI

57.6

N/A

36

Claude Opus 4.1

Anthropic

57.3

N/A

37
M

Muse Spark

Meta

57.0

N/A

38
B

ERNIE 5.0

Baidu

56.8

N/A

39

DeepSeek-V4-Flash-Max

DeepSeek

56.7

$0.1 in / $0.2 out

40

Claude Opus 4.7

Anthropic

56.5

$5 in / $25 out