Grok 4.5
Comprehensive benchmark performance across 11 evaluation categories
Composite Score
0.0/100Rank
Token Benchmark
10.7Lower burn, higher score
Total Tokens
~83.1k/test
Category Radar
The full radar chart is shown on wider screens. On mobile, the category breakdown below provides the same values in a readable stacked layout.
Historical Composite Score
Category Breakdown
| # | Category | Score | Tests | 7-Day Trend | Weight |
|---|---|---|---|---|---|
| 1 | Bug Introduction Rate | 96.3 | 3 | 1.0x | |
| 2 | Coding Tasks | 96.0 | 3 | 1.0x | |
| 3 | Bug Fixes | 92.7 | 3 | 1.0x | |
| 4 | Performance & Efficiency | 90.0 | 3 | 1.0x | |
| 5 | Security Awareness | 86.9 | 3 | 1.0x | |
| 6 | Instruction Following | 83.3 | 3 | 1.0x | |
| 7 | Feature Implementation | 82.7 | 3 | 1.0x | |
| 8 | Code Thoroughness | 77.6 | 3 | 1.0x | |
| 9 | Code Quality | 73.7 | 3 | 1.0x | |
| 10 | Long Reasoning | 46.4 | 3 | 1.0x | |
| 11 | Token Efficiency | 10.7 | 30 | 1.0x |
Individual Test Results
Token Efficiency is computed from every successful task in the run. The model with the lowest average token burn receives 100, and heavier token usage is penalized proportionally.
Avg Tokens/Test
83.1k
Total Tokens
2.5M
Regression History
21Score dropped -9.2% from 89.2 to 81.0
Score dropped -23.8% from 83.6 to 63.7
Score dropped -10.1% from 85.6 to 77.0
Score dropped -12.2% from 87.7 to 77.0
Score dropped -8.7% from 74.1 to 67.7
Score dropped -25.8% from 90.8 to 67.3
Score dropped -8.8% from 92.9 to 84.7
Score dropped -8.6% from 79.6 to 72.7
Score dropped -3.3% from 86.8 to 83.9
Score dropped -3.8% from 92.9 to 89.4
Score dropped -3.2% from 97.8 to 94.7
Score dropped -5.4% from 73.8 to 69.8
Score dropped -3.3% from 96.4 to 93.3
Score dropped -8.3% from 93.5 to 85.7
Score dropped -6.0% from 93.3 to 87.7
Score dropped -4.0% from 98.2 to 94.3
Score dropped -3.3% from 93.3 to 90.2
Score dropped -5.5% from 87.2 to 82.4
Score dropped -31.6% from 98.4 to 67.3
Score dropped -30.7% from 97.1 to 67.3
Score dropped -78.5% from 61.5 to 13.2