Skytells
HomeModelsCLIChangelog
  • Home
  • Models
  • CLI
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Documentation
  • API Reference
  • Pricing
  • ModelsNew

Platform

  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network
  • Trust Center
  • CLI

Resources

  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

334

Tracked models

29

Providers

286

Benchmarked

13.1

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

334 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
41

Step-3.5-Flash

step-3.5-flash

codeprogrammingtool use
SStepFun

59.8

Inference

62.859.836.548.693.9$0.1 in / $0.4 out
42

Gemini 3.1 Pro

gemini-3.1-pro-preview

multimodalvisionmulti-input reasoning
Google

58.3

Inference

72.358.352.660.922.7
43

Claude Haiku 4.5

claude-haiku-4-5-20251001

multimodalvisionmulti-input reasoning
Anthropic

53.3

Inference

30.853.350.853.945.6
44

DeepSeek-V3.2 (Non-thinking)

deepseek-chat

textinference
DeepSeek

52.0

Inference

0.052.00.00.082.7$0.28 in / $0.42 out
45

Gemini 2.5 Pro

gemini-2.5-pro

multimodalvisionmulti-input reasoning
Google

51.0

Inference

42.551.00.021.429.8
46

GPT-4 Turbo

gpt-4-turbo-2024-04-09

textinference
OpenAI

50.2

Inference

15.550.20.00.015.4$10 in / $30 out
47

GPT-5.3 Chat

gpt-5.3-chat-latest

multimodalvisionmulti-input reasoning
OpenAI

50.2

Inference

0.050.20.00.031.5
48

Grok-3

grok-3

multimodalvisionmulti-input reasoning
xAI

50.2

Inference

58.450.20.00.024.1$3 in / $15 out
49

GPT-5.1 Medium

gpt-5.1-medium-2025-11-12

multimodalvisionmulti-input reasoning
OpenAI

44.9

Inference

64.144.90.00.032.0
50

GPT-5.4 Mini

gpt-5.4-mini

texttext-to-textlanguage
OpenAI

44.5

Inference

51.144.515.020.042.7
51

GPT-5.4 nano

gpt-5.4-nano

multimodalvisionmulti-input reasoning
OpenAI

44.5

Inference

41.844.56.98.276.8$0.2 in / $1.25 out
52

GPT-4o

gpt-4o-2024-08-06

multimodalvisionmulti-input reasoning
OpenAI

39.6

Inference

29.139.614.93.731.2
53

Grok 4.5

grok-4.5

multimodalvisionmulti-input reasoning
xAI

38.2

Inference

69.338.20.070.835.6$2 in / $6 out
54

GPT-4o

gpt-4o-2024-05-13

multimodalvisionmulti-input reasoning
OpenAI

38.0

Inference

20.538.00.00.030.7
55

GPT-5.4

gpt-5.4

texttext-to-textlanguage
OpenAI

37.2

Inference

70.537.248.950.318.5
56

Gemma 4 26B-A4B

gemma-4-26b-a4b-it

multimodalvisionmulti-input reasoning
Google

32.9

Inference

43.832.90.00.090.2
57

Gemma 4 31B

gemma-4-31b-it

multimodalvisionmulti-input reasoning
Google

32.9

Inference

55.432.90.00.091.5
58

Kimi K2.6

kimi-k2.6

texttext-to-textlanguage
Moonshot AI

32.9

Inference

63.532.947.269.446.3
59

Kimi K2.7 Code

kimi-k2.7-code

multimodalvisionmulti-input reasoning
Moonshot AI

32.9

Inference

0.032.947.20.047.6
60

Nemotron 3 Nano (30B A3B)

nemotron-3-nano-30b-a3b

codeprogrammingtool use
NNVIDIA

32.9

Inference

43.632.93.03.8100.0$0.06 in / $0.24 out
41
S

Step-3.5-Flash

StepFun

59.8

$0.1 in / $0.4 out

42

Gemini 3.1 Pro

Google

58.3

$2.5 in / $15 out

43

Claude Haiku 4.5

Anthropic

53.3

$1 in / $5 out

44

Page 3 of 17 · 334 models

PreviousNext

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$2.5 in / $15 out
$1 in / $5 out
$1.25 in / $10 out
$1.75 in / $14 out
$1.25 in / $10 out
$0.75 in / $4.5 out
$2.5 in / $10 out
$2.5 in / $10 out
$2.5 in / $15 out
$0.13 in / $0.4 out
$0.13 in / $0.38 out
$0.75 in / $3.5 out
$0.74 in / $3.5 out

DeepSeek-V3.2 (Non-thinking)

DeepSeek

52.0

$0.28 in / $0.42 out

45

Gemini 2.5 Pro

Google

51.0

$1.25 in / $10 out

46

GPT-4 Turbo

OpenAI

50.2

$10 in / $30 out

47

GPT-5.3 Chat

OpenAI

50.2

$1.75 in / $14 out

48

Grok-3

xAI

50.2

$3 in / $15 out

49

GPT-5.1 Medium

OpenAI

44.9

$1.25 in / $10 out

50

GPT-5.4 Mini

OpenAI

44.5

$0.75 in / $4.5 out

51

GPT-5.4 nano

OpenAI

44.5

$0.2 in / $1.25 out

52

GPT-4o

OpenAI

39.6

$2.5 in / $10 out

53

Grok 4.5

xAI

38.2

$2 in / $6 out

54

GPT-4o

OpenAI

38.0

$2.5 in / $10 out

55

GPT-5.4

OpenAI

37.2

$2.5 in / $15 out

56

Gemma 4 26B-A4B

Google

32.9

$0.13 in / $0.4 out

57

Gemma 4 31B

Google

32.9

$0.13 in / $0.38 out

58

Kimi K2.6

Moonshot AI

32.9

$0.75 in / $3.5 out

59

Kimi K2.7 Code

Moonshot AI

32.9

$0.74 in / $3.5 out

60
N

Nemotron 3 Nano (30B A3B)

NVIDIA

32.9

$0.06 in / $0.24 out