Back to Dashboard
CategoryWeight: 1.0x

Bug Introduction Rate

Measures how often the model introduces new bugs while writing or modifying code. Lower is better (inverted for scoring).

Best Score

0.0

Avg Score

0.0

Tests

3

Performance Over Time — All Models

Model Rankings

1
Grok 4.5

Category score

View
96.3BEST
Tokens159.1k
Total159.1k
2
Claude Opus 5

Category score

View
95.3-1.0 pts
Tokens17.6k
Total17.6k
3
GPT-5.6 Sol

Category score

View
91.0-5.3 pts
Tokens51.0k
Total51.0k
4
Claude Sonnet 5

Category score

View
90.7-5.6 pts
Tokens15.0k
Total15.0k

Test Breakdown

Refactor Without Regression

Refactor a function without introducing new failures in existing tests

Grok 4.5
95.0
Claude Opus 5
96.0
GPT-5.6 Sol
88.0
Claude Sonnet 5
96.0

Merge Conflict Resolution

Resolve merge conflicts without introducing semantic errors

Grok 4.5
97.0
Claude Opus 5
97.0
GPT-5.6 Sol
97.0
Claude Sonnet 5
96.0

Dependency Upgrade Safety

Upgrade a dependency and adapt code without breaking changes

Grok 4.5
97.0
Claude Opus 5
93.0
GPT-5.6 Sol
88.0
Claude Sonnet 5
80.0