Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

15.1

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
61

DeepSeek-V3.2-Speciale

deepseek-v3.2-speciale

codeprogrammingtool use
DeepSeek

42.0

Programming

50.70.05.042.00.0N/A
62

Kimi K2.5

kimi-k2.5

multimodalvisionmulti-input reasoning
Moonshot AI

41.8

Programming

63.70.041.041.80.0N/A
63

DeepSeek-V4-Flash-Max

deepseek-v4-flash-max

codeprogrammingtool use
DeepSeek

41.4

Programming

56.284.835.341.498.8
64

Nemotron 3 Ultra (550B A55B)

nemotron-3-ultra-550b-a55b

codeprogrammingtool use
NNVIDIA

40.6

Programming

54.70.011.540.60.0N/A
65

MiniMax M2

minimax-m2

codeprogrammingtool use
MiniMax

40.0

Programming

29.862.840.240.073.2$0.3 in / $1.2 out
66

Qwen3.6-27B

qwen3.6-27b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

39.4

Programming

54.732.90.039.449.4$0.6 in / $3.6 out
67

Qwen3.5-27B

qwen3.5-27b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

39.1

Programming

57.832.941.339.161.0$0.3 in / $2.4 out
68

Qwen3.5-122B-A10B

qwen3.5-122b-a10b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

38.3

Programming

60.50.044.638.30.0N/A
69

DeepSeek-V3.2-Exp

deepseek-v3.2-exp

codeprogrammingtool use
DeepSeek

38.0

Programming

50.20.027.238.00.0N/A
70

Claude 3.7 Sonnet

claude-3-7-sonnet-20250219

multimodalvisionmulti-input reasoning
Anthropic

37.7

Programming

42.30.049.137.70.0
71

LongCat-Flash-Chat

longcat-flash-chat

codeprogrammingtool use
Meituan

36.6

Programming

26.00.048.136.60.0N/A
72

Muse Spark

muse-spark

multimodalvisionmulti-input reasoning
MMeta

36.2

Programming

67.10.064.436.20.0N/A
73

GLM-4.5

glm-4.5

codeprogrammingtool use
ZZhipu AI

36.1

Programming

31.50.036.036.10.0N/A
74

Grok Code Fast 1

grok-code-fast-1

codeprogrammingtool use
xAI

36.1

Programming

0.027.20.036.160.5$0.2 in / $1.5 out
75

MiMo-V2-Flash

mimo-v2-flash

codeprogrammingtool use
Xiaomi

35.8

Programming

50.60.023.535.80.0N/A
76

GPT-5.3 Codex

gpt-5.3-codex

texttext-to-textcoding
OpenAI

34.5

Programming

0.031.10.034.522.0
77

LongCat-Flash-Thinking-2601

longcat-flash-thinking-2601

codeprogrammingtool use
Meituan

33.5

Programming

52.90.025.633.50.0
78

MAI-Thinking-1

mai-thinking-1

codeprogrammingtool use
MMicrosoft

32.2

Programming

60.10.00.032.20.0N/A
79

Qwen3-Coder 480B A35B Instruct

qwen3-coder-480b-a35b-instruct

codeprogrammingtool use
AAlibaba Cloud / Qwen Team

32.1

Programming

0.00.050.732.10.0
80

Qwen3 Max

qwen3-max

codeprogrammingtool use
AAlibaba Cloud / Qwen Team

32.1

Programming

28.00.00.032.10.0N/A
61

DeepSeek-V3.2-Speciale

DeepSeek

42.0

N/A

62

Kimi K2.5

Moonshot AI

41.8

N/A

63

DeepSeek-V4-Flash-Max

DeepSeek

41.4

$0.1 in / $0.2 out

64

Page 4 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$0.1 in / $0.2 out
N/A
$1.75 in / $14 out
N/A
N/A
N

Nemotron 3 Ultra (550B A55B)

NVIDIA

40.6

N/A

65

MiniMax M2

MiniMax

40.0

$0.3 in / $1.2 out

66
A

Qwen3.6-27B

Alibaba Cloud / Qwen Team

39.4

$0.6 in / $3.6 out

67
A

Qwen3.5-27B

Alibaba Cloud / Qwen Team

39.1

$0.3 in / $2.4 out

68
A

Qwen3.5-122B-A10B

Alibaba Cloud / Qwen Team

38.3

N/A

69

DeepSeek-V3.2-Exp

DeepSeek

38.0

N/A

70

Claude 3.7 Sonnet

Anthropic

37.7

N/A

71

LongCat-Flash-Chat

Meituan

36.6

N/A

72
M

Muse Spark

Meta

36.2

N/A

73
Z

GLM-4.5

Zhipu AI

36.1

N/A

74

Grok Code Fast 1

xAI

36.1

$0.2 in / $1.5 out

75

MiMo-V2-Flash

Xiaomi

35.8

N/A

76

GPT-5.3 Codex

OpenAI

34.5

$1.75 in / $14 out

77

LongCat-Flash-Thinking-2601

Meituan

33.5

N/A

78
M

MAI-Thinking-1

Microsoft

32.2

N/A

79
A

Qwen3-Coder 480B A35B Instruct

Alibaba Cloud / Qwen Team

32.1

N/A

80
A

Qwen3 Max

Alibaba Cloud / Qwen Team

32.1

N/A