Anthropic also publishes a rung above Opus: Claude Fable 5 at $10 in and $50 out per 1M tokens, twice the Opus tier and ten times Haiku. Counting it, the full published spread across Claude's own ladder is 10×.
What TSK-1 Found
TSK-1 has not run Opus and Sonnet head-to-head in the same build: they are tiers of the same family, not competing models. What the tests do show is that Opus produced the best-looking build we have measured (Jul 30, 2026) and kept the fullest record of its own work (Aug 1, 2026), while Sonnet writes the cleanest code we have measured (Aug 3, 2026) and is the only model that checked its own agent by chatting with it and confirming the answers came from the app's own data (Jul 30, 2026). Sonnet is the better all-purpose choice; Opus is the quality ceiling.
- Opus: Jul 30, 2026 — the best-looking build we have measured; Aug 1, 2026 — the fullest record of its own work.
- Sonnet: Aug 3, 2026 — the cleanest code we have measured; Jul 30, 2026 — checked its own agent by chatting with it and confirming the answers came from the app's own data.
See the full evidence at /tsk/claude and the TSK-1 hub.
The Headline
Anthropic ships Claude across a price ladder, and the wrong rung on the wrong task is the most expensive mistake in 2026 AI workflows. Run the Opus tier on a triage step that Haiku handles cleanly and you pay five times more for the same outcome. Run Fable on it and you pay ten times more. Run Haiku on a graduate-reasoning task and you get a wrong answer cheaper.
Here is the part most comparisons miss: that 10× internal spread is wider than the quality gap between most frontier labs. Picking the right rung inside Claude usually moves your bill more than switching vendors does. Tier discipline beats vendor shopping.
The right rule: start at Sonnet, move up to Opus only for the hardest 10%, drop down to Haiku for the high-volume 30%.
TL;DR: Sonnet is the workhorse default at $2 in / $10 out per 1M tokens (introductory through 31 August 2026, then $3 / $15). Opus is for the hardest 10% at $5 / $25. Haiku is for triage, routing, and bulk steps at $1 / $5. Fable 5 sits above at $10 / $50. Inside Taskade Genesis every rung lives in the same picker with cost shown per tier in the tooltip.
The Claude Tier Ladder
All prices above are Anthropic's published US-dollar list rates per million tokens, read on 2026-08-11. Two adjustments apply on top: batch requests are 50% off in both directions, and a cache hit bills at a tenth of base input. A cache-heavy agent loop on Sonnet can therefore land well under the headline rate.
Two Pricing Facts That Change the Arithmetic
Before you route anything, two published Anthropic behaviours matter more than a few cents on the rate card.
1. The full 1M-token context bills at standard rates. A 900,000-token request costs the same per token as a 9,000-token one. There is no long-context surcharge and no threshold where the rate steps up. That makes whole-repository analysis and multi-document synthesis genuinely budgetable on the Opus and Sonnet tiers — you can size the job from token count alone.
2. Claude 4.7 and later tokenize about 30% denser. The newer tokenizer produces roughly 30% more tokens for the same text. So price per token is not price per page. A rate-card win of 20% against another model family can vanish entirely once the same document is tokenized. This cuts both ways inside Claude too: comparing a 4.5-era measurement against a 4.7-era one will flatter the older model.
The practical rule that falls out of both: benchmark on a representative document from your own corpus and compare billed totals, not headline rates. It takes an afternoon and it is the only method that survives a tokenizer change.
When to Pick Each: A Practical Decision Tree
The Tier-Stacking Pattern (Cuts Cost Without Hurting Quality)
The most effective Claude pattern in 2026 is not picking one tier. It is stacking all three across a single workflow so each step runs on the cheapest tier that gets it right.
Workflow: Customer support escalation
┌────────────────────────────────────────────────────────────┐
│ STEP 1: Classify incoming ticket │
│ → Haiku ($5 per 1M output — a fifth of the Opus tier) │
│ │
│ STEP 2: Retrieve customer context, extract fields │
│ → Haiku ($5 per 1M output) │
│ │
│ STEP 3: Draft response with product knowledge │
│ → Sonnet ($10 per 1M output — the workhorse) │
│ │
│ STEP 4: Review for tone and compliance │
│ → Sonnet ($10 per 1M output) │
│ │
│ STEP 5: Escalation cases only, re-draft with nuance │
│ → Opus ($25 per 1M output, on ~10% of tickets) │
└────────────────────────────────────────────────────────────┘
The arithmetic, assuming five steps of roughly equal output volume
and escalation on 10% of tickets (output-token rates, per 1M):
Opus tier for every step: 5 × $25 = $125
Tier-stacked as above: 2×$5 + 2×$10 + 0.1×$25 = $32.50
≈ 74% less, with quality unchanged on the 10% that matters
Those assumptions are stated on purpose. Change the step count or the escalation rate and the saving moves — but the shape holds, because the cheapest and most expensive rungs you are mixing differ by 5×.
Inside Taskade Genesis each agent or automation step can pick a different tier from the model picker. Build the workflow once. Pick tiers per step. The credit math takes care of itself.
Opus vs Sonnet on the Most Common Workloads
Direct head-to-head on workloads where teams typically face the choice.
| Workload | Opus | Sonnet | Reach for | Cost consequence (out /1M) |
|---|---|---|---|---|
| Conversational chat agent | excellent | excellent | Sonnet — indistinguishable quality | $10 vs $25 |
| Code completion / pair programming | excellent | excellent | Sonnet — best cost-to-quality | $10 vs $25 |
| Agentic code-edit loops | strong | strong | Sonnet, cache-heavy | $10, or ~$1 on cache hits |
| Long-form blog post drafting | strongest | strong | Opus if brand quality matters | $25 — 2.5× the Sonnet rate |
| Customer-facing email reply | strong | strong | Sonnet for most, Opus for VIP | $10 baseline, $25 on the exceptions |
| Graduate-level scientific reasoning | strongest | competitive | Opus — genuine quality gap | $25 |
| Math reasoning (AIME, HMMT) | strongest | competitive | Opus for the hardest | $25 |
| Multilingual content | strong | strong | Sonnet unless the language is rare | $10 |
| Multi-document research synthesis | strongest | strong | Opus — and the 1M window carries no surcharge | $25 flat across the window |
| Peak-nuance single answers | — | — | Fable 5, sparingly | $50 — 5× Sonnet |
| Support classification / routing | overkill | overkill | Haiku — skip both | $5 |
| Bulk data extraction | overkill | overkill | Haiku | $5 |
| On-prem or self-hosted | ✗ | ✗ | neither — Claude is hosted-API only | open-weight models are the only route |
Note the pattern. Sonnet is the right answer in most rows. Opus has a real quality edge in about three categories. That edge is worth roughly 2.5× the token cost only when the task genuinely needs it — and Haiku, at a fifth of the Opus rate, quietly handles more rows than most teams expect.
Where Sonnet Loses to Opus
Be honest about it. Three categories where Opus's edge is real and visible.
- Truly hard reasoning. Multi-hop puzzles, mathematical proofs, novel scientific reasoning. Opus does not just score higher on a benchmark, it gets to the answer Sonnet cannot. At 2.5× the token cost this is the clearest-cut upgrade on the ladder.
- Long-form prose where nuance matters. Brand-critical writing, executive communication, sensitive customer email. Opus's prose has a polish Sonnet matches 95% of the time but misses on the hard cases.
- Safety-critical reasoning. When the cost of a wrong answer is high (medical, legal, financial advice contexts), Opus's Constitutional AI training shows up more clearly in edge cases. This is also when human review remains mandatory.
Outside these three, Sonnet is the right default.
The Taskade Genesis Angle: All Three, Plus Open-Source
The smartest 2026 pattern is not picking among Opus, Sonnet, and Haiku. It is mixing all three with open-source picks across the same workflow.
Inside Taskade Genesis the model picker shows credit cost per option. Auto mode handles tier selection if you do not want to think about it. You can override on any step. Taskade routes across 15+ frontier models from OpenAI, Anthropic, Google, and open-weight providers inside one workspace, with the AI allowance included in the subscription rather than a separate consumer subscription per vendor — so you are not betting the company on one lab's roadmap or one lab's next price change.
Three patterns that work well right now.
- Pattern 1: Tier-stacked Claude. Haiku for triage. Sonnet for the bulk of the work. Opus for the final answer. On the five-step illustration above that is roughly 74% less than running the Opus tier end to end.
- Pattern 2: Sonnet + open-weight. Sonnet for the chat surface where polished conversation matters. An open-weight coding model for agentic work inside the same workspace. DeepSeek V4 Pro for bulk extraction — it lists at $1.32 in and $3.96 out per 1M at peak, halving off-peak outside 01:00–04:00 and 06:00–10:00 UTC on weekdays, so scheduling matters as much as model choice.
- Pattern 3: Opus only for the moments that matter. Default everything to Sonnet or an open-weight model. Reserve the Opus tier for steps where the customer or the legal team will read the output. Spend the savings on more iterations.
See 10 Best Open-Source AI LLMs in 2026 for how the open-source picks map onto the Claude tier ladder.
Final Word: The Tier Discipline
The biggest cost mistake teams make with Claude in 2026 is running the Opus tier by default. Switch your default to Sonnet at $2 in and $10 out per 1M. Use Haiku at $1 and $5 for the triage steps Sonnet does not need to run. Reserve Opus at $5 and $25 for the 10% of tasks where the quality genuinely justifies the step up, and treat Fable 5 at $10 and $50 as an exception, not a setting.
Then diarise two dates. 1 September 2026, when Sonnet 5's introductory rate ends and it moves to $3 and $15. And whenever you next change model versions, because the 4.7-and-later tokenizer emits about 30% more tokens for the same text — a version bump can raise your bill without a single price changing.
Inside Taskade Genesis you do not have to hold all of that in your head. The model picker shows the cost per option. Auto mode picks for you. The savings compound.
▲ Memory feeds Intelligence. ■ Intelligence triggers Execution. ● Execution creates Memory. A full Claude ladder. Open-weight brains alongside it. One workspace. The right model for every step.
This is the origin of living software. 🌱
Build with Opus, Sonnet, and Haiku in one workspace →
Related reading
- 10 Best Open-Source AI LLMs in 2026 — The full ranking and where each fits.
- Multi-Model AI Access — How Taskade Genesis routes 15+ models.
- Kimi vs Claude — Open-source agentic coding vs premium frontier chat.
- Qwen vs DeepSeek — The two open-source frontier giants.
- Free Claude Alternative — Genesis as a workspace alternative.
- Tools for AI Agents — The built-in agent toolkit.
- TSK-1 Claude profile — Full benchmark evidence for the Claude family.
- TSK-1 hub — The complete model benchmark dataset.
