Back to Dashboard
CategoryWeight: 1.0x

Code Quality

Evaluates readability, idiomatic patterns, naming conventions, and adherence to language best practices.

Best Score

0.0

Avg Score

0.0

Tests

3

Performance Over Time — All Models

Model Rankings

1
GPT-5.6 Sol

Category score

View
92.0BEST
Tokens90.7k
Total90.7k
2
Claude Sonnet 5

Category score

View
89.7-2.3 pts
Tokens88.9k
Total88.9k
3
Claude Opus 5

Category score

View
87.7-4.3 pts
Tokens101.3k
Total101.3k
4
Grok 4.5

Category score

View
73.7-18.3 pts
Tokens469.2k
Total469.2k

Test Breakdown

Idiomatic Python

Write Pythonic code using generators, comprehensions, and context managers

GPT-5.6 Sol
95.0
Claude Sonnet 5
96.0
Claude Opus 5
94.0
Grok 4.5
78.0

TypeScript Best Practices

Use strict types, discriminated unions, and proper error narrowing

GPT-5.6 Sol
90.0
Claude Sonnet 5
91.0
Claude Opus 5
90.0
Grok 4.5
65.0

Clean Architecture Patterns

Implement repository pattern with proper dependency inversion

GPT-5.6 Sol
91.0
Claude Sonnet 5
82.0
Claude Opus 5
79.0
Grok 4.5
78.0