Claude Sonnet 5 Benchmarks: Full Score Breakdown vs Opus 4.8, Sonnet 4.6, GPT-5.5 & More
Anthropic launched Claude Sonnet 5 on June 30, 2026, and along with it came one of the densest benchmark releases the company has published for a Sonnet-class model. Beyond the headline "near-Opus performance, lower price" pitch, the actual numbers tell a more specific story — where Sonnet 5 genuinely closes the gap with Opus 4.8, where it still falls short, and where it surprisingly pulls ahead.
This article breaks down every benchmark Anthropic and independent testers have published so far, with direct comparisons to Sonnet 4.6, Opus 4.8, and rival models like GPT-5.5, Gemini, and GLM-5.2.
Quick summary
- Sonnet 5 improves on every single published benchmark compared to Sonnet 4.6
- It closes most, but not all, of the gap to Opus 4.8 — Opus still leads on the hardest agentic coding and reasoning tasks
- On one benchmark (GDPval-AA v2, knowledge work), Sonnet 5 actually edges past Opus 4.8
- It beats GPT-5.5 on every directly comparable benchmark while costing significantly less
- Cybersecurity capability remains deliberately weak — both Sonnet 5 and Sonnet 4.6 scored 0% on a Firefox exploit-development test
Coding benchmarks
Coding is where Sonnet 5 makes its strongest case, and where Anthropic published the most data.
| Benchmark | Sonnet 5 | Sonnet 4.6 | Opus 4.8 | Gain over 4.6 |
|---|---|---|---|---|
| SWE-bench Verified | 85.2% | — | — | — |
| SWE-bench Pro | 63.2% | 58.1% | 69.2% | +5.1 |
| SWE-bench Multilingual | 78.3% | — | — | — |
| SWE-bench Multimodal | 28.1% | — | — | — |
| Terminal-Bench 2.1 | 80.4% | 67.0% | — | +13.4 |
| FrontierCode v1 | 38.8% | 15.1% | — | +23.7 |
SWE-bench Pro is the headline number — it's a harder, less-leaked benchmark than the original SWE-bench, testing larger multi-file diffs. Sonnet 5's jump from 58.1% to 63.2% is a meaningful improvement, though Opus 4.8 still leads by 6 points.
Terminal-Bench 2.1 shows the biggest practical jump — a 13.4-point gain over Sonnet 4.6. This matters directly for anyone running CLI-based agents, repo-maintenance bots, or Claude Code workflows, since this benchmark measures performance inside terminal and command-line environments specifically.
FrontierCode v1 more than doubled — from 15.1% to 38.8%. This benchmark tests novel algorithmic problems that can't be solved by pattern-matching against training data, so the gain reflects a structural improvement in reasoning rather than just familiarity with common coding patterns.
According to Anthropic's system card, both Sonnet models use the same evaluation harness (mini-SWE-agent), making the comparison a clean apples-to-apples test. The practical read: tasks Sonnet 4.6 used to fail around a third of the time are now succeeding roughly 80% of the time under Sonnet 5.
Read Now: Claude AI Token Pricing Risks: A 2026 Guide to Controlling Your AI Costs
Agentic and computer-use benchmarks
These benchmarks measure how well the model can plan, navigate interfaces, and complete multi-step tasks autonomously — the core selling point of this release.
| Benchmark | Sonnet 5 | Sonnet 4.6 | Opus 4.8 |
|---|---|---|---|
| OSWorld-Verified (computer use) | 81.2% | 78.5% | — |
| BrowseComp (single agent) | 84.7% | — | — |
| BrowseComp (multi-agent) | 86.6% | — | — |
| Humanity's Last Exam (with tools) | 57.4% | — | 57.9% |
| AutomationBench | 13.5% | 5.3% | — |
AutomationBench is worth calling out specifically because the absolute numbers look low. This benchmark requires an agent to complete a full cross-app business workflow end to end, with correct API calls at every step — one mistake anywhere in the chain fails the whole task. Going from 5.3% to 13.5% is a large relative improvement (2.5x), even though the pass rate is still modest in absolute terms. That's the nature of a strict, all-or-nothing benchmark.
Humanity's Last Exam (HLE) with tools is one of the closest results on the entire scorecard — Sonnet 5 hits 57.4% against Opus 4.8's 57.9%, a gap of just half a point. This is one of the hardest general reasoning benchmarks available, so landing this close to Opus here is a genuinely strong result for a mid-tier model.
Knowledge work and document understanding
| Benchmark | Sonnet 5 (no tools) | Sonnet 5 (with tools) | Opus 4.8 |
|---|---|---|---|
| GDPval-AA v2 (knowledge work) | — | 1,618 Elo | 1,615 Elo |
| GDP.pdf (document understanding) | 67.5% | 81.6% | — |
| ChartMuseum | 70.1% | 86.7% | — |
| CharXiv Reasoning | 77.0% | 88.3% | — |
| BenchCAD Vision2Code | 0.266 IoU | 0.373 IoU | — |
GDPval-AA v2 is the one benchmark where Sonnet 5 actually beats Opus 4.8 — 1,618 Elo versus 1,615. This benchmark measures real knowledge-work tasks, and the result is close enough to be statistically tied, but it's still notable that a cheaper mid-tier model can match or edge out the flagship here.
Read Now: Claude AI Token Pricing Risks: A 2026 Guide to Controlling Your AI Costs
The document and chart understanding benchmarks (GDP.pdf, ChartMuseum, CharXiv) all show a consistent pattern: Sonnet 5 performs solidly without tools, but improves substantially once it's allowed to use tools (like code execution to parse a chart or re-render a document). If your workflow involves PDFs, screenshots, charts, or CAD-style visual reasoning, testing with tools enabled is where you'll see the real capability gain.
Cybersecurity benchmarks
Anthropic deliberately kept Sonnet 5's cyber capability low, and the numbers back that up.
| Test | Sonnet 5 | Sonnet 4.6 | Opus 4.8 | Mythos 5 |
|---|---|---|---|---|
| Firefox 147 exploit dev (working exploit) | 0.0% | 0.0% | 68.8% | 88.4% |
| Firefox 147 exploit dev (partial success) | 13.2% | 8.8% | — | — |
Neither Sonnet model could produce a single working exploit in this test, built in collaboration with Mozilla. Sonnet 5 did show a slightly higher partial-success rate than Sonnet 4.6, but Anthropic attributes this to general intelligence gains rather than any deliberate cybersecurity training. Cyber safeguards ship enabled by default on Sonnet 5 — the same real-time detection system used on Opus 4.7 and 4.8.
Sonnet 5 vs GPT-5.5
Anthropic and third-party benchmarking sites report Sonnet 5 beating GPT-5.5 on every directly comparable metric:
| Benchmark | Sonnet 5 | GPT-5.5 | Sonnet 5 advantage |
|---|---|---|---|
| SWE-bench Pro | 63.2% | 58.6% | +4.6 |
| Terminal-Bench 2.1 | 80.4% | 78.2% | +2.2 |
| HLE (with tools) | 57.4% | 52.2% | +5.2 |
On pricing, Sonnet 5 launched at $2/$10 per million tokens (rising to $3/$15 after August 31) against GPT-5.5's $5/$30 — meaning Sonnet 5 currently wins on both capability and cost in this matchup.
Sonnet 5 vs Gemini 3.5 Flash
This comparison is more of a trade-off than a clean win. Sonnet 5 leads on coding benchmarks (+8.1 on SWE-bench Pro, +4.2 on Terminal-Bench), but Gemini 3.5 Flash leads on MCP Atlas (83.6%), runs roughly 4x faster (289 tokens/sec), and costs about half as much. If raw coding capability is the priority, Sonnet 5 wins; if throughput and cost efficiency matter more, Gemini 3.5 Flash is the stronger pick.
Sonnet 5 vs GLM-5.2
On one provisional third-party leaderboard, Sonnet 5 leads the overall aggregate score 94 to 90 against GLM-5.2. The breakdown:
- Coding: Sonnet 5 ahead, 63.2 vs 62.1 average
- Agentic tasks: Sonnet 5 ahead, 81.8 vs 81.0 average
- Knowledge: GLM-5.2 ahead, 67.2 vs 57.4 average
- HLE specifically: Sonnet 5 leads sharply, 57.4% vs 54.7%
The pattern here mirrors the Opus 4.8 comparison — Sonnet 5 is the stronger all-round agentic and coding model, but pure knowledge-recall tasks are where competitors can still pull ahead.
What the benchmarks don't tell you
Worth flagging: Anthropic issued a correction on launch day, noting that its original cost-performance chart for the BrowseComp evaluation used a simpler methodology than the one it standardly applies to agentic search evaluations, and republished it with the standard method. It's a reminder that benchmark charts, even from the model maker itself, get revised — worth checking you're looking at the latest version before citing a number.
It's also worth noting that Sonnet 5 uses an updated tokenizer (the same one introduced with Opus 4.7). The same input text can map to roughly 1.0–1.35x more tokens than on previous models — something to account for when comparing real-world cost, not just the headline per-token price.
Conclusion
Across nearly every published benchmark, Claude Sonnet 5 is a clear, consistent step up from Sonnet 4.6 — the gains on Terminal-Bench (+13.4), FrontierCode (+23.7 points, more than doubling), and AutomationBench (2.5x improvement) are the standout numbers. Against Opus 4.8, the gap has narrowed substantially and, on knowledge-work tasks, has effectively closed. Against outside competitors like GPT-5.5, Sonnet 5 wins comfortably on both benchmarks and price. The honest takeaway: it's not the most capable model Anthropic makes, but on a benchmark-per-dollar basis, it's currently the strongest option in the Claude lineup for most agentic and coding work.
Frequently Asked Questions
What is Sonnet 5's score on SWE-bench?
Sonnet 5 scores 85.2% on SWE-bench Verified and 63.2% on the harder SWE-bench Pro, up from Sonnet 4.6's 58.1% on SWE-bench Pro.
Does Sonnet 5 beat Opus 4.8 on any benchmark?
Yes — on the GDPval-AA v2 knowledge-work benchmark, Sonnet 5 scores 1,618 Elo against Opus 4.8's 1,615, effectively a statistical tie with Sonnet 5 slightly ahead. On most other benchmarks, Opus 4.8 remains ahead.
How much better is Sonnet 5 than Sonnet 4.6?
On every published benchmark, Sonnet 5 improves over Sonnet 4.6. The largest gains are on FrontierCode v1 (15.1% to 38.8%), Terminal-Bench 2.1 (67.0% to 80.4%), and AutomationBench (5.3% to 13.5%).
Is Sonnet 5 good for coding?
Yes, coding is Sonnet 5's strongest area. It scores 85.2% on SWE-bench Verified and shows the biggest jump of any category on Terminal-Bench 2.1, making it well suited for CLI-based agents and Claude Code workflows.
Does Sonnet 5 have strong cybersecurity capability?
No, deliberately so. Both Sonnet 5 and Sonnet 4.6 scored 0% on a Firefox exploit-development test built with Mozilla, far below Opus 4.8 (68.8%) and Mythos 5 (88.4%). Anthropic recommends Opus 4.8 for sanctioned security research work.
Does Claude Sonnet 5 think for longer than Opus?
They are tuned differently rather than one simply being slower. Sonnet 5 is built for throughput and cost efficiency, so on comparable tasks it generally produces an answer with less deliberation than an Opus-class model, which is the trade being made. If your workload depends on extended reasoning rather than speed, benchmark both on your own prompts — published scores do not capture how much thinking time each model spends on your particular task.
Is Claude Sonnet 5 good enough to replace Opus?
For most day-to-day coding, drafting and document work the benchmark gap is small enough that Sonnet 5 is the sensible default on cost grounds. The gap widens on long multi-step agentic tasks and the hardest reasoning problems, which is where an Opus-class model still earns its price.



1 Comment
The comparison between Sonnet 5 and Opus 4.8 was especially useful because it highlights where the smaller model is genuinely competitive instead of just focusing on overall scores. I’d also be interested in seeing more independent real-world evaluations over longer workflows, since benchmark gains don’t always translate equally into day-to-day coding or research tasks.
Community guidelines
We want the comments to be useful for every reader. Every comment is reviewed before it is published. A comment will not be approved if it is:
Keep it genuine and on-topic and it will be approved quickly. Thank you for helping keep the discussion clean.