Back to Dashboard
OpenAI

GPT-5.6 Sol

Comprehensive benchmark performance across 11 evaluation categories

Composite Score

0.0/100

Rank

#3

Token Benchmark

45.4

Lower burn, higher score

Total Tokens

588.4k

~19.6k/test

Category Radar

The full radar chart is shown on wider screens. On mobile, the category breakdown below provides the same values in a readable stacked layout.

Historical Composite Score

Category Breakdown

3 tests1.0x weight
3 tests1.0x weight
3 tests1.0x weight
3 tests1.0x weight
3 tests1.0x weight
3 tests1.0x weight
3 tests1.0x weight
88.0
3 tests1.0x weight
3 tests1.0x weight
30 tests1.0x weight
3 tests1.0x weight

Individual Test Results

Token Efficiency is computed from every successful task in the run. The model with the lowest average token burn receives 100, and heavier token usage is penalized proportionally.

Avg Tokens/Test

19.6k

Total Tokens

588.4k

Mathematical Proof
15.3k tok23.3s77.5
Legal Reasoning Chain
15.7k tok26.0s25.0
Multi-step Logic Puzzle
16.5k tok51.3s17.0
Graph Algorithm Implementation
15.2k tok22.7s92.0
REST API Design
14.6k tok9.0s96.0
Concurrent Data Pipeline
14.9k tok13.1s90.0
Off-by-One Boundary Fix
15.2k tok20.4s74.0
Race Condition Detection
14.7k tok15.8s100.0
Memory Leak Fix
15.2k tok24.7s90.0
OAuth2 Integration
15.9k tok31.8s97.0
Search Autocomplete
15.8k tok30.0s97.0
Webhook System
15.8k tok27.1s88.0
Test Suite Completeness
17.1k tok49.9s88.0
Edge Case Coverage
31.6k tok5m 50s90.0
Error Path Completeness
64.4k tok6m 23s94.6
Refactor Without Regression
15.7k tok19.5s88.0
Merge Conflict Resolution
16.1k tok24.2s97.0
Dependency Upgrade Safety
19.2k tok1m 17s88.0
SQL Injection Prevention
15.9k tok33.4s70.0
Secret Management
16.4k tok57.9s89.0
XSS Mitigation
21.5k tok3m 21s100.0
Structured Output Compliance
14.6k tok10.8s100.0
Constraint Adherence
14.5k tok6.2s100.0
Multi-step Instruction Chain
14.4k tok11.8s100.0
TypeScript Best Practices
18.5k tok1m 32s90.0
Idiomatic Python
54.0k tok2m 16s95.0
Clean Architecture Patterns
18.3k tok1m 19s91.0
Algorithm Complexity
15.6k tok21.9s97.0
Memory-efficient Processing
17.6k tok1m 20s91.0
Query Optimization
18.3k tok1m 39s88.0

Regression History

11
Code Qualityminor

Score dropped -3.9% from 89.1 to 85.7

Detected Aug 6, 2026
Feature Implementationminorresolved

Score dropped -4.3% from 90.6 to 86.7

Detected Aug 5, 2026·Resolved Aug 9, 2026
Coding Tasksminor

Score dropped -3.3% from 93.8 to 90.7

Detected Aug 5, 2026
Token Efficiencymoderate

Score dropped -7.8% from 44.0 to 40.6

Detected Aug 3, 2026
Long Reasoningmoderate

Score dropped -14.1% from 42.5 to 36.5

Detected Aug 3, 2026
Coding Tasksminorresolved

Score dropped -3.5% from 94.3 to 91.0

Detected Aug 2, 2026·Resolved Aug 2, 2026
Code Thoroughnessmajorresolved

Score dropped -18.9% from 90.7 to 73.5

Detected Aug 2, 2026·Resolved Aug 2, 2026
Bug Introduction Ratemajor

Score dropped -35.3% from 91.1 to 59.0

Detected Jul 31, 2026
Security Awarenessmoderate

Score dropped -5.7% from 84.0 to 79.2

Detected Jul 28, 2026
Code Qualityminorresolved

Score dropped -4.8% from 90.0 to 85.7

Detected Jul 28, 2026·Resolved Jul 28, 2026
Bug Fixesminorresolved

Score dropped -4.8% from 84.4 to 80.4

Detected Jul 28, 2026·Resolved Jul 28, 2026

Outage History

10
error

Started Aug 10, 3:00 AM·Ended Aug 10, 3:00 AM· checks affected
error

Started Aug 8, 3:00 AM·Ended Aug 8, 3:00 AM· checks affected
error

Started Aug 7, 3:00 AM·Ended Aug 7, 3:00 AM· checks affected
error

Started Aug 6, 3:00 AM·Ended Aug 6, 3:00 AM· checks affected
error

Started Aug 5, 3:00 AM·Ended Aug 5, 3:00 AM· checks affected
error

Started Aug 4, 3:00 AM·Ended Aug 4, 3:00 AM· checks affected
error

Started Aug 2, 3:00 AM·Ended Aug 2, 3:00 AM· checks affected
error

Started Jul 31, 3:00 AM·Ended Jul 31, 3:00 AM· checks affected
error

Started Jul 29, 3:00 AM·Ended Jul 29, 3:00 AM· checks affected
error

Started Jul 26, 3:00 AM·Ended Jul 26, 3:00 AM· checks affected