Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

12.2

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
21

Gemini 3.1 Pro

gemini-3.1-pro-preview

multimodalvisionmulti-input reasoning
Google

52.6

Agentic

72.358.352.660.922.7$2.5 in / $15 out
22

DeepSeek-V4-Pro-Max

deepseek-v4-pro-max

codeprogrammingtool use
DeepSeek

51.0

Agentic

64.984.851.052.045.1
23

Qwen3 VL 235B A22B Instruct

qwen3-vl-235b-a22b-instruct

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

51.0

Agentic

34.30.051.00.00.0
24

Claude Haiku 4.5

claude-haiku-4-5-20251001

multimodalvisionmulti-input reasoning
Anthropic

50.8

Agentic

30.853.350.853.945.6
25

Kimi K2-Thinking-0905

kimi-k2-thinking-0905

codeprogrammingtool use
Moonshot AI

50.8

Agentic

66.20.050.858.00.0
26

Qwen3-Coder 480B A35B Instruct

qwen3-coder-480b-a35b-instruct

codeprogrammingtool use
AAlibaba Cloud / Qwen Team

50.7

Agentic

0.00.050.732.10.0
27

Claude Opus 4.6

claude-opus-4-6

multimodalvisionmulti-input reasoning
Anthropic

49.8

Agentic

74.628.349.871.58.0
28

Qwen3.7 Max

qwen3.7-max

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

49.6

Agentic

64.762.849.674.243.9$1.25 in / $3.75 out
29

Claude Sonnet 4

claude-sonnet-4-20250514

multimodalvisionmulti-input reasoning
Anthropic

49.4

Agentic

39.30.049.442.80.0
30

Claude 3.7 Sonnet

claude-3-7-sonnet-20250219

multimodalvisionmulti-input reasoning
Anthropic

49.1

Agentic

42.30.049.137.70.0
31

GLM-5V-Turbo

glm-5v-turbo

multimodalvisionmulti-input reasoning
ZZhipu AI

49.1

Agentic

0.00.049.10.00.0N/A
32

GPT-5.4

gpt-5.4

texttext-to-textlanguage
OpenAI

48.9

Agentic

70.537.248.950.318.5
33

LongCat-Flash-Chat

longcat-flash-chat

codeprogrammingtool use
Meituan

48.1

Agentic

26.00.048.136.60.0N/A
34

Qwen3.7-Plus

qwen3.7-plus

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

48.1

Agentic

62.962.848.161.169.5$0.32 in / $1.28 out
35

Kimi K2.6

kimi-k2.6

texttext-to-textlanguage
Moonshot AI

47.2

Agentic

63.532.947.269.446.3
36

Kimi K2.7 Code

kimi-k2.7-code

multimodalvisionmulti-input reasoning
Moonshot AI

47.2

Agentic

0.032.947.20.047.6
37

GPT-5.2 Pro

gpt-5.2-pro-2025-12-11

multimodalvisionmulti-input reasoning
OpenAI

46.1

Agentic

64.40.046.10.00.0
38

MiniMax M2.1

minimax-m2.1

codeprogrammingtool use
MiniMax

45.7

Agentic

39.168.945.747.672.9$0.3 in / $1.2 out
39

Seed 2.0 Pro

seed-2.0-pro

multimodalvisionmulti-input reasoning
BByteDance

44.8

Agentic

66.823.244.856.054.9$0.5 in / $3 out
40

o1

o1-2024-12-17

multimodalvisionmulti-input reasoning
OpenAI

44.7

Agentic

41.90.044.75.60.0N/A
21

Gemini 3.1 Pro

Google

52.6

$2.5 in / $15 out

22

DeepSeek-V4-Pro-Max

DeepSeek

51.0

$1.6 in / $3.2 out

23
A

Qwen3 VL 235B A22B Instruct

Alibaba Cloud / Qwen Team

51.0

N/A

24

Page 2 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$1.6 in / $3.2 out
N/A
$1 in / $5 out
N/A
N/A
$5 in / $25 out
N/A
N/A
$2.5 in / $15 out
$0.75 in / $3.5 out
$0.74 in / $3.5 out
N/A

Claude Haiku 4.5

Anthropic

50.8

$1 in / $5 out

25

Kimi K2-Thinking-0905

Moonshot AI

50.8

N/A

26
A

Qwen3-Coder 480B A35B Instruct

Alibaba Cloud / Qwen Team

50.7

N/A

27

Claude Opus 4.6

Anthropic

49.8

$5 in / $25 out

28
A

Qwen3.7 Max

Alibaba Cloud / Qwen Team

49.6

$1.25 in / $3.75 out

29

Claude Sonnet 4

Anthropic

49.4

N/A

30

Claude 3.7 Sonnet

Anthropic

49.1

N/A

31
Z

GLM-5V-Turbo

Zhipu AI

49.1

N/A

32

GPT-5.4

OpenAI

48.9

$2.5 in / $15 out

33

LongCat-Flash-Chat

Meituan

48.1

N/A

34
A

Qwen3.7-Plus

Alibaba Cloud / Qwen Team

48.1

$0.32 in / $1.28 out

35

Kimi K2.6

Moonshot AI

47.2

$0.75 in / $3.5 out

36

Kimi K2.7 Code

Moonshot AI

47.2

$0.74 in / $3.5 out

37

GPT-5.2 Pro

OpenAI

46.1

N/A

38

MiniMax M2.1

MiniMax

45.7

$0.3 in / $1.2 out

39
B

Seed 2.0 Pro

ByteDance

44.8

$0.5 in / $3 out

40

o1

OpenAI

44.7

N/A