Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

15.1

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
21

Gemini 3.1 Pro

gemini-3.1-pro-preview

multimodalvisionmulti-input reasoning
Google

60.9

Programming

72.358.352.660.922.7$2.5 in / $15 out
22

Claude Opus 4.1

claude-opus-4-1-20250805

multimodalvisionmulti-input reasoning
Anthropic

60.8

Programming

46.00.067.460.80.0
23

GLM-5

glm-5

codeprogrammingtool use
ZZhipu AI

60.4

Programming

0.06.935.160.440.0$1 in / $3.2 out
24

Seed 2.1 Pro

seed-2.1-pro

multimodalvisionmulti-input reasoning
BByteDance

60.2

Programming

69.20.075.660.20.0N/A
25

GLM-5.2

glm-5.2

codeprogrammingtool use
ZZhipu AI

59.9

Programming

68.784.844.159.951.2$0.95 in / $3 out
26

Hy3

hy3

codeprogrammingtool use
TTencent

59.2

Programming

65.90.040.459.20.0N/A
27

Mistral Medium 3.5

mistral-medium-3-5

multimodalvisionmulti-input reasoning
Mistral AI

59.0

Programming

34.623.215.459.034.8
28

Muse Spark 1.1

muse-spark-1.1

multimodalvisionmulti-input reasoning
MMeta

58.1

Programming

69.784.876.658.141.5$1.25 in / $4.25 out
29

Kimi K2-Thinking-0905

kimi-k2-thinking-0905

codeprogrammingtool use
Moonshot AI

58.0

Programming

66.20.050.858.00.0
30

MiMo-V2.5-Pro

mimo-v2.5-pro

codeprogrammingtool use
Xiaomi

56.3

Programming

36.284.80.056.378.0$0.435 in / $0.87 out
31

Seed 2.0 Pro

seed-2.0-pro

multimodalvisionmulti-input reasoning
BByteDance

56.0

Programming

66.823.244.856.054.9$0.5 in / $3 out
32

Qwen3.5-397B-A17B

qwen3.5-397b-a17b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

55.3

Programming

54.50.024.755.30.0N/A
33

Claude Haiku 4.5

claude-haiku-4-5-20251001

multimodalvisionmulti-input reasoning
Anthropic

53.9

Programming

30.853.350.853.945.6
34

Seed 2.1 Turbo

seed-2.1-turbo

multimodalvisionmulti-input reasoning
BByteDance

53.8

Programming

66.30.063.153.80.0N/A
35

Qwen3.6 Plus

qwen3.6-plus

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

52.8

Programming

66.962.833.652.854.9$0.5 in / $3 out
36

Gemini 3.6 Flash

gemini-3.6-flash

multimodalvisionmulti-input reasoning
Google

52.6

Programming

57.884.80.052.634.8
37

Gemini 3 Pro

gemini-3-pro-preview

multimodalvisionmulti-input reasoning
Google

52.3

Programming

70.30.058.052.30.0
38

DeepSeek-V4-Pro-Max

deepseek-v4-pro-max

codeprogrammingtool use
DeepSeek

52.0

Programming

64.984.851.052.045.1
39

GPT-5.1

gpt-5.1-2025-11-13

multimodalvisionmulti-input reasoning
OpenAI

51.4

Programming

62.462.80.051.437.3
40

GPT-5.1 Instant

gpt-5.1-instant-2025-11-12

multimodalvisionmulti-input reasoning
OpenAI

51.4

Programming

62.462.80.051.437.3
21

Gemini 3.1 Pro

Google

60.9

$2.5 in / $15 out

22

Claude Opus 4.1

Anthropic

60.8

N/A

23
Z

GLM-5

Zhipu AI

60.4

$1 in / $3.2 out

24

Page 2 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

N/A
$1.5 in / $7.5 out
N/A
$1 in / $5 out
$1.5 in / $7.5 out
N/A
$1.6 in / $3.2 out
$1.25 in / $10 out
$1.25 in / $10 out
B

Seed 2.1 Pro

ByteDance

60.2

N/A

25
Z

GLM-5.2

Zhipu AI

59.9

$0.95 in / $3 out

26
T

Hy3

Tencent

59.2

N/A

27

Mistral Medium 3.5

Mistral AI

59.0

$1.5 in / $7.5 out

28
M

Muse Spark 1.1

Meta

58.1

$1.25 in / $4.25 out

29

Kimi K2-Thinking-0905

Moonshot AI

58.0

N/A

30

MiMo-V2.5-Pro

Xiaomi

56.3

$0.435 in / $0.87 out

31
B

Seed 2.0 Pro

ByteDance

56.0

$0.5 in / $3 out

32
A

Qwen3.5-397B-A17B

Alibaba Cloud / Qwen Team

55.3

N/A

33

Claude Haiku 4.5

Anthropic

53.9

$1 in / $5 out

34
B

Seed 2.1 Turbo

ByteDance

53.8

N/A

35
A

Qwen3.6 Plus

Alibaba Cloud / Qwen Team

52.8

$0.5 in / $3 out

36

Gemini 3.6 Flash

Google

52.6

$1.5 in / $7.5 out

37

Gemini 3 Pro

Google

52.3

N/A

38

DeepSeek-V4-Pro-Max

DeepSeek

52.0

$1.6 in / $3.2 out

39

GPT-5.1

OpenAI

51.4

$1.25 in / $10 out

40

GPT-5.1 Instant

OpenAI

51.4

$1.25 in / $10 out