Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

12.2

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
61

Claude Opus 4.5

claude-opus-4-5-20251101

multimodalvisionmulti-input reasoning
Anthropic

35.2

Agentic

54.50.035.272.20.0N/A
62

GLM-5.1

glm-5.1

codeprogrammingtool use
ZZhipu AI

35.2

Agentic

62.615.935.247.240.2$1.4 in / $4.4 out
63

GLM-5

glm-5

codeprogrammingtool use
ZZhipu AI

35.1

Agentic

0.06.935.160.440.0$1 in / $3.2 out
64

Qwen3.6 Plus

qwen3.6-plus

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

33.6

Agentic

66.962.833.652.854.9$0.5 in / $3 out
65

Gemini 3 Flash

gemini-3-flash-preview

multimodalvisionmulti-input reasoning
Google

33.2

Agentic

68.562.833.261.954.9
66

GPT-4.1

gpt-4.1-2025-04-14

multimodalvisionmulti-input reasoning
OpenAI

32.8

Agentic

27.273.232.814.740.7
67

Qwen3 VL 32B Thinking

qwen3-vl-32b-thinking

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

31.1

Agentic

41.20.031.10.00.0
68

LongCat-Flash-Lite

longcat-flash-lite

codeprogrammingtool use
Meituan

30.1

Agentic

22.872.830.123.995.6$0.1 in / $0.4 out
69

DeepSeek-V3.2-Exp

deepseek-v3.2-exp

codeprogrammingtool use
DeepSeek

27.2

Agentic

50.20.027.238.00.0N/A
70

GPT OSS 120B

gpt-oss-120b

textinference
OpenAI

26.8

Agentic

33.70.026.80.00.0N/A
71

MiniMax M1 40K

minimax-m1-40k

codeprogrammingtool use
MiniMax

26.8

Agentic

21.30.026.815.50.0N/A
72

Qwen3-235B-A22B-Thinking-2507

qwen3-235b-a22b-thinking-2507

textinference
AAlibaba Cloud / Qwen Team

26.8

Agentic

44.40.026.80.00.0N/A
73

GLM-4.7

glm-4.7

multimodalvisionmulti-input reasoning
ZZhipu AI

26.6

Agentic

60.30.026.642.50.0N/A
74

MiniMax M2.7

minimax-m2.7

codeprogrammingtool use
MiniMax

26.3

Agentic

0.019.526.329.073.2$0.3 in / $1.2 out
75

LongCat-Flash-Thinking-2601

longcat-flash-thinking-2601

codeprogrammingtool use
Meituan

25.6

Agentic

52.90.025.633.50.0
76

Qwen3 VL 32B Instruct

qwen3-vl-32b-instruct

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

25.1

Agentic

26.50.025.10.00.0
77

Qwen3.5-397B-A17B

qwen3.5-397b-a17b

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

24.7

Agentic

54.50.024.755.30.0N/A
78

GPT-5

gpt-5-2025-08-07

multimodalvisionmulti-input reasoning
OpenAI

24.3

Agentic

59.90.024.347.10.0N/A
79

GLM-4.5-Air

glm-4.5-air

codeprogrammingtool use
ZZhipu AI

24.2

Agentic

25.70.024.216.00.0N/A
80

Qwen3 VL 8B Instruct

qwen3-vl-8b-instruct

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

24.0

Agentic

8.00.024.00.00.0N/A
61

Claude Opus 4.5

Anthropic

35.2

N/A

62
Z

GLM-5.1

Zhipu AI

35.2

$1.4 in / $4.4 out

63
Z

GLM-5

Zhipu AI

35.1

$1 in / $3.2 out

64

Page 4 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$0.5 in / $3 out
$2 in / $8 out
N/A
N/A
N/A
A

Qwen3.6 Plus

Alibaba Cloud / Qwen Team

33.6

$0.5 in / $3 out

65

Gemini 3 Flash

Google

33.2

$0.5 in / $3 out

66

GPT-4.1

OpenAI

32.8

$2 in / $8 out

67
A

Qwen3 VL 32B Thinking

Alibaba Cloud / Qwen Team

31.1

N/A

68

LongCat-Flash-Lite

Meituan

30.1

$0.1 in / $0.4 out

69

DeepSeek-V3.2-Exp

DeepSeek

27.2

N/A

70

GPT OSS 120B

OpenAI

26.8

N/A

71

MiniMax M1 40K

MiniMax

26.8

N/A

72
A

Qwen3-235B-A22B-Thinking-2507

Alibaba Cloud / Qwen Team

26.8

N/A

73
Z

GLM-4.7

Zhipu AI

26.6

N/A

74

MiniMax M2.7

MiniMax

26.3

$0.3 in / $1.2 out

75

LongCat-Flash-Thinking-2601

Meituan

25.6

N/A

76
A

Qwen3 VL 32B Instruct

Alibaba Cloud / Qwen Team

25.1

N/A

77
A

Qwen3.5-397B-A17B

Alibaba Cloud / Qwen Team

24.7

N/A

78

GPT-5

OpenAI

24.3

N/A

79
Z

GLM-4.5-Air

Zhipu AI

24.2

N/A

80
A

Qwen3 VL 8B Instruct

Alibaba Cloud / Qwen Team

24.0

N/A