Back to Dashboard
CategoryWeight: 1.0x

Instruction Following

Measures how closely the model adheres to explicit constraints in the prompt, including formatting, language, and output structure.

Best Score

0.0

Avg Score

0.0

Tests

3

Performance Over Time — All Models

Model Rankings

1
GPT-5.6 Sol

Category score

View
100.0BEST
Tokens43.5k
Total43.5k
2
Claude Opus 5

Category score

View
94.7-5.3 pts
Tokens859
Total859
3
Grok 4.5

Category score

View
83.3-16.7 pts
Tokens57.6k
Total57.6k
4
Claude Sonnet 5

Category score

View
82.7-17.3 pts
Tokens1.0k
Total1.0k

Test Breakdown

Structured Output Compliance

Produce JSON matching an exact schema with no extra fields

GPT-5.6 Sol
100.0
Claude Opus 5
100.0
Grok 4.5
60.0
Claude Sonnet 5
100.0

Constraint Adherence

Follow explicit constraints like max line length and naming conventions

GPT-5.6 Sol
100.0
Claude Opus 5
94.0
Grok 4.5
100.0
Claude Sonnet 5
73.0

Multi-step Instruction Chain

Execute a 6-step instruction sequence without skipping or reordering

GPT-5.6 Sol
100.0
Claude Opus 5
90.0
Grok 4.5
90.0
Claude Sonnet 5
75.0