Back to Dashboard
CategoryWeight: 1.0x

Token Efficiency

Measures how efficiently a model solves tasks by penalizing higher token consumption. Lower usage earns a higher score.

Best Score

0.0

Avg Score

0.0

Measured Models

4

Performance Over Time — All Models

Model Rankings

1
Claude Sonnet 5

Token usage benchmark

View
100.0BEST
Avg/Test8.9k/test
Total267.2k
2
Claude Opus 5

Token usage benchmark

View
66.7-33.3 pts
Avg/Test13.4k/test
Total400.6k
3
GPT-5.6 Sol

Token usage benchmark

View
45.4-54.6 pts
Avg/Test19.6k/test
Total588.4k
4
Grok 4.5

Token usage benchmark

View
10.7-89.3 pts
Avg/Test83.1k/test
Total2.5M

Benchmark Construction

How It Scores

We total prompt and completion tokens across all successful benchmark tasks, compute an average per successful task, then assign 100 to the lowest-burn model in that run. Everyone else is scaled down proportionally, so higher usage means a lower benchmark score.

What To Read In This View

The ranking table above is the benchmark itself. Use Avg/Test to compare per-task burn and Total to spot larger absolute usage across the whole run.

Historical scores show whether a model is becoming more or less token-efficient over time, independent of raw quality improvements in the other ten categories.