GPT-5.6 Sol
Comprehensive benchmark performance across 11 evaluation categories
Composite Score
0.0/100Rank
Token Benchmark
45.4Lower burn, higher score
Total Tokens
~19.6k/test
Category Radar
The full radar chart is shown on wider screens. On mobile, the category breakdown below provides the same values in a readable stacked layout.
Historical Composite Score
Category Breakdown
| # | Category | Score | Tests | 7-Day Trend | Weight |
|---|---|---|---|---|---|
| 1 | Instruction Following | 100.0 | 3 | 1.0x | |
| 2 | Feature Implementation | 94.0 | 3 | 1.0x | |
| 3 | Coding Tasks | 92.7 | 3 | 1.0x | |
| 4 | Code Quality | 92.0 | 3 | 1.0x | |
| 5 | Performance & Efficiency | 92.0 | 3 | 1.0x | |
| 6 | Bug Introduction Rate | 91.0 | 3 | 1.0x | |
| 7 | Code Thoroughness | 90.9 | 3 | 1.0x | |
| 8 | Bug Fixes | 88.0 | 3 | 1.0x | |
| 9 | Security Awareness | 86.3 | 3 | 1.0x | |
| 10 | Token Efficiency | 45.4 | 30 | 1.0x | |
| 11 | Long Reasoning | 39.8 | 3 | 1.0x |
Individual Test Results
Token Efficiency is computed from every successful task in the run. The model with the lowest average token burn receives 100, and heavier token usage is penalized proportionally.
Avg Tokens/Test
19.6k
Total Tokens
588.4k
Regression History
11Score dropped -3.9% from 89.1 to 85.7
Score dropped -4.3% from 90.6 to 86.7
Score dropped -3.3% from 93.8 to 90.7
Score dropped -7.8% from 44.0 to 40.6
Score dropped -14.1% from 42.5 to 36.5
Score dropped -3.5% from 94.3 to 91.0
Score dropped -18.9% from 90.7 to 73.5
Score dropped -35.3% from 91.1 to 59.0
Score dropped -5.7% from 84.0 to 79.2
Score dropped -4.8% from 90.0 to 85.7
Score dropped -4.8% from 84.4 to 80.4