Skytells
  • Home
  • Changelog
Skytells

Addressing the world's greatest challenges with AI. Enterprise research, foundation models, and infrastructure trusted by organizations worldwide since 2012.

Get Started

  • Console
  • Learn
  • Pricing
  • ModelsNew
  • EveNew
  • CLI

Platform

  • DropsNew
  • Cognition
  • Cloud AgentsNew
  • AI Solutions
  • Infrastructure
  • Edge Network

Resources

  • Documentation
  • Blog
  • Changelog
  • AI Leaderboard
  • Research
  • Status

Company

  • About
  • Careers
  • Trust Center
  • Legal
  • Privacy Policy

© 2012–2026 Skytells, Inc. All rights reserved.

Live rankings

AI Model Leaderboard

Every major AI model ranked across benchmark quality, inference speed, agentic capability, programming aptitude, and cost efficiency — updated continuously from published evaluation data.

Explore full leaderboardBrowse model catalog

373

Tracked models

34

Providers

310

Benchmarked

14.1

Avg. index

OverallBenchmarksInferenceAgenticProgrammingValue / Price

373 models

RankModelProviderScoreBenchmarksInferenceAgenticProgrammingValuePrice
1

Nemotron 3.5 Lightning (30B A3B)

nemotron-3.5-lightning-30b-a3b

codeprogrammingtool use
NNVIDIA

100.0

Value / Price

26.630.56.99.0100.0$0.05 in / $0.2 out
2

DeepSeek-V4-Flash-0731

deepseek-v4-flash-0731

textinference
DeepSeek

99.0

Value / Price

0.081.956.30.099.0$0.09 in / $0.18 out
3

Nemotron 3 Nano (30B A3B)

nemotron-3-nano-30b-a3b

codeprogrammingtool use
NNVIDIA

98.1

Value / Price

42.430.53.05.998.1
4

DeepSeek-V4-Flash-0423

deepseek-v4-flash-0423

codeprogrammingtool use
DeepSeek

95.1

Value / Price

51.481.923.240.695.1
5

Laguna S 2.1

laguna-s-2.1

codeprogrammingtool use
PPoolside

95.1

Value / Price

0.081.932.752.295.1$0.1 in / $0.2 out
6

Laguna XS 2.1

laguna-xs-2.1

codeprogrammingtool use
PPoolside

95.1

Value / Price

0.030.50.022.395.1$0.1 in / $0.2 out
7

Muse Spark 1.2

muse-spark-1.2

multimodalvisionmulti-input reasoning
MMeta

95.1

Value / Price

0.081.90.00.095.1$0.1 in / $0.2 out
8

Muse Spark 1.3

muse-spark-1.3

multimodalvisionmulti-input reasoning
MMeta

95.1

Value / Price

70.481.90.00.095.1$0.1 in / $0.2 out
9

DeepSeek-V4-Flash-Max

deepseek-v4-flash-max

codeprogrammingtool use
DeepSeek

91.3

Value / Price

55.381.932.244.191.3
10

LongCat-Flash-Lite

longcat-flash-lite

codeprogrammingtool use
Meituan

91.0

Value / Price

22.371.930.124.191.0
11

GPT-4.1 nano

gpt-4.1-nano-2025-04-14

multimodalvisionmulti-input reasoning
OpenAI

90.3

Value / Price

11.185.90.00.090.3
12

Step-3.5-Flash

step-3.5-flash

codeprogrammingtool use
SStepFun

89.3

Value / Price

63.260.236.548.889.3$0.1 in / $0.4 out
13

MiMo-V2.5

mimo-v2.5

multimodalvisionmulti-input reasoning
Xiaomi

87.4

Value / Price

46.581.90.029.787.4$0.168 in / $0.336 out
14

Gemma 4 31B

gemma-4-31b-it

multimodalvisionmulti-input reasoning
Google

86.4

Value / Price

52.930.50.00.086.4
15

Gemma 4 26B-A4B

gemma-4-26b-a4b-it

multimodalvisionmulti-input reasoning
Google

85.4

Value / Price

41.630.50.00.085.4
16

Qwen3.8 Flash

qwen3.8-flash

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

83.5

Value / Price

61.158.664.560.083.5$0.15 in / $0.47 out
17

GLM-5.3-Flash

glm-5.3-flash

multimodalvisionmulti-input reasoning
ZZhipu AI

82.5

Value / Price

64.081.972.60.082.5$0.15 in / $0.5 out
18

Mercury 2

mercury-2

codeprogrammingtool use
IInception

79.8

Value / Price

42.168.70.021.079.8$0.25 in / $0.75 out
19

Qwen3 VL 4B Instruct

qwen3-vl-4b-instruct

multimodalvisionmulti-input reasoning
AAlibaba Cloud / Qwen Team

79.6

Value / Price

18.330.516.10.079.6
20

DeepSeek-V3.2 (Non-thinking)

deepseek-chat

textinference
DeepSeek

77.3

Value / Price

0.048.70.00.077.3$0.28 in / $0.42 out
1
N

Nemotron 3.5 Lightning (30B A3B)

NVIDIA

100.0

$0.05 in / $0.2 out

2

DeepSeek-V4-Flash-0731

DeepSeek

99.0

$0.09 in / $0.18 out

3
N

Nemotron 3 Nano (30B A3B)

NVIDIA

98.1

$0.06 in / $0.24 out

Page 1 of 19 · 373 models

Next

Want benchmark charts, model comparison, and pricing analytics?

Sign in to access the full interactive leaderboard with deep benchmark breakdowns and model comparison tools.

Open full leaderboard

Rankings are based on multi-dimensional evaluation across benchmark quality, inference efficiency, and cost-per-output. Scores are updated continuously and may differ from individual third-party benchmarks.

$0.06 in / $0.24 out
$0.1 in / $0.2 out
$0.14 in / $0.28 out
$0.1 in / $0.4 out
$0.1 in / $0.4 out
$0.13 in / $0.38 out
$0.13 in / $0.4 out
$0.1 in / $0.6 out
4

DeepSeek-V4-Flash-0423

DeepSeek

95.1

$0.1 in / $0.2 out

5
P

Laguna S 2.1

Poolside

95.1

$0.1 in / $0.2 out

6
P

Laguna XS 2.1

Poolside

95.1

$0.1 in / $0.2 out

7
M

Muse Spark 1.2

Meta

95.1

$0.1 in / $0.2 out

8
M

Muse Spark 1.3

Meta

95.1

$0.1 in / $0.2 out

9

DeepSeek-V4-Flash-Max

DeepSeek

91.3

$0.14 in / $0.28 out

10

LongCat-Flash-Lite

Meituan

91.0

$0.1 in / $0.4 out

11

GPT-4.1 nano

OpenAI

90.3

$0.1 in / $0.4 out

12
S

Step-3.5-Flash

StepFun

89.3

$0.1 in / $0.4 out

13

MiMo-V2.5

Xiaomi

87.4

$0.168 in / $0.336 out

14

Gemma 4 31B

Google

86.4

$0.13 in / $0.38 out

15

Gemma 4 26B-A4B

Google

85.4

$0.13 in / $0.4 out

16
A

Qwen3.8 Flash

Alibaba Cloud / Qwen Team

83.5

$0.15 in / $0.47 out

17
Z

GLM-5.3-Flash

Zhipu AI

82.5

$0.15 in / $0.5 out

18
I

Mercury 2

Inception

79.8

$0.25 in / $0.75 out

19
A

Qwen3 VL 4B Instruct

Alibaba Cloud / Qwen Team

79.6

$0.1 in / $0.6 out

20

DeepSeek-V3.2 (Non-thinking)

DeepSeek

77.3

$0.28 in / $0.42 out