Best LLM for Coding 2026 | AI Coding Model Rankings & Benchmarks | Onyx AI
Coding LLM Leaderboard
Which AI model writes the best code? We rank every major LLM — open and closed source — across software engineering, code generation, competitive programming, and agentic coding benchmarks.
Roshan Desai · Last updated: 2026-07-20
CodingReasoning
| Model | Rank | Params | Cost per 1M | Performance % | Best Use Case |
|---|---|---|---|---|---|
| Claude Fable 5 | S | - | $10.00 | 100% | Software Engineering |
| GPT-5.6 Sol | S | - | $5.00 | 95% | Software Engineering |
| Claude Opus 4.8 | A | - | $5.00 | 88.6% | Software Engineering |
| GPT-5.5 | A | - | $5.00 | 88.7% | Software Engineering |
| Claude Sonnet 5 | A | - | $2.00 | 85.2% | Software Engineering |
| Kimi K3 | S | 2.8T | $3.00 | 80% | Code Generation |
| MiniMax M3 | S | 428B | $0.30 | 80.5% | Code Generation |
| DeepSeek-V4-Pro | A | 1.6T | $0.43 | 80.6% | Code Generation |
Cost vs. Coding Performance
Which models give you the best coding performance for the price? Top-left is the sweet spot — high performance, low cost.
Best Coding LLMs by Benchmark
Best at Software Engineering
How does each model perform on real-world software engineering tasks? Hover any bar for details.
| Model | Performance % |
|---|---|
| Claude Fable 5 | 100% |
| GPT-5.5 | 88.7% |
| Claude Opus 4.8 | 88.6% |
| Claude Sonnet 5 | 85.2% |
| MiniMax M3 | 80.5% |
Best for Code Generation
| Model | Performance % |
|---|---|
| Grok 3 | 100% |
| GPT-5.5 | 85% |
| Gemini 3.1 Pro | 93% |
Best in Competitive Coding
| Model | Performance % |
|---|---|
| DeepSeek-V4-Pro | 100% |
| GPT-5.5 | 85% |
| DeepSeek R1 | 49% |
Best at Terminal Coding
| Model | Performance % |
|---|---|
| GPT-5.6 Sol | 89% |
| Claude Fable 5 | 85% |
| Kimi K3 | 81% |
Hardest Software Engineering Tasks
| Model | Performance % |
|---|---|
| Claude Fable 5 | 50% |
| GPT-5.6 Sol | 68% |
| Claude Opus 4.8 | 59% |
Coding Benchmark Scores & Pricing
Complete coding benchmark results and pricing for every model. Click any column header to sort.
| Model | Params | Input $/1M | Output $/1M | SWE-bench Verified | LiveCodeBench | Terminal-Bench 2.0 |
|---|---|---|---|---|---|---|
| Claude Fable 5 | - | $10.00 | $50.00 | 95.0 | N/A | 84.3 |
| Claude Opus 4.8 | - | $5.00 | $25.00 | 88.6 | N/A | 74.6 |
| Claude Sonnet 5 | - | $2.00 | $10.00 | 85.2 | N/A | 80.4 |
| DeepSeek R1 | 671B | $0.28 | $0.42 | 49.2 | 90.2 | N/A |
| DeepSeek V3.2 | 685B | $0.28 | $0.42 | 67.8 | N/A | N/A |
| DeepSeek-V4-Flash | 284B | $0.14 | $0.28 | 79.0 | N/A | N/A |
| DeepSeek-V4-Pro | 1.6T | $0.43 | $0.87 | 80.6 | N/A | N/A |
Compare Coding LLMs Head-to-Head
Select two models to see how they compare across all coding and reasoning benchmarks.
Comparison: Claude Opus 4.8 vs GPT-5.5
| Benchmark | Claude Opus 4.8 | GPT-5.5 |
|---|---|---|
| GPQA Diamond | 93.6 | 93.6 |
| SWE-bench Verified | 88.6 | 88.7 |
| SWE-bench Pro | 69.2 | 59.4 |
| Terminal-Bench 2.1 | 74.6 | 85.6 |
| Benchmarks won | 1 | 2 |