Back to Dashboard
CategoryWeight: 1.0x

Code Thoroughness

Evaluates completeness of generated code: edge case handling, input validation, error paths, and test coverage.

Best Score

0.0

Avg Score

0.0

Tests

3

Performance Over Time — All Models

Model Rankings

1
GPT-5.6 Sol

Category score

View
90.9BEST
Tokens113.1k
Total113.1k
2
Claude Sonnet 5

Category score

View
86.2-4.7 pts
Tokens73.8k
Total73.8k
3
Claude Opus 5

Category score

View
84.5-6.4 pts
Tokens134.1k
Total134.1k
4
Grok 4.5

Category score

View
77.6-13.3 pts
Tokens1.1M
Total1.1M

Test Breakdown

Edge Case Coverage

Generate code handling null, empty, unicode, and overflow inputs

GPT-5.6 Sol
90.0
Claude Sonnet 5
82.0
Claude Opus 5
96.0
Grok 4.5
88.0

Error Path Completeness

Ensure all failure modes have proper error handling and logging

GPT-5.6 Sol
94.6
Claude Sonnet 5
94.6
Claude Opus 5
94.6
Grok 4.5
82.8

Test Suite Completeness

Generate tests covering happy path, edge cases, and integration

GPT-5.6 Sol
88.0
Claude Sonnet 5
82.0
Claude Opus 5
63.0
Grok 4.5
62.0