Across Haiku 4.5, Sonnet 5, Opus 5, and Fable 5.1, output price spans $5 to $50 per million tokens, unchanged since a Sonnet 5 hike to $3 and $15 was cancelled before its September 1, 2026 start date.
Haiku 4.5 wins on price. Stacking caching with batch cuts a typical bill by 65%. Read the caching section first.
Model | Input, $/MTok | Output, $/MTok | 5-min cache write | 1-hr cache write | Cache read | Batch input / output | Min. cacheable tokens |
|---|---|---|---|---|---|---|---|
Claude Haiku 4.5 | $1.00 | $5.00 | $1.25 | $2.00 | $0.10 | $0.50 / $2.50 | 4,096 |
Claude Sonnet 5 | $2.00 | $10.00 | $2.50 | $4.00 | $0.20 | $1.00 / $5.00 | 1,024 |
Claude Opus 5 | $5.00 | $25.00 | $6.25 | $10.00 | $0.50 | $2.50 / $12.50 | 512 |
Claude Fable 5.1 | $10.00 | $50.00 | $12.50 | $20.00 | $0.25 † | $5.00 / $25.00 | 512 |
Table note: † Claude Fable 5.1 and Claude Mythos 5.1 price cache reads at 0.025 times the standard input rate rather than the usual 0.1 times; every other model in this table reads a cache hit at 0.1 times its own input price. All rates are direct-API rates with inference_geo left at its global default.
LLM Waves Research received no pre-release access, free credits, or rate-limit exemption from Anthropic for this piece, and Anthropic did not see a draft before publication. This is a Basis 2 (aggregated) article: every rate below is collected from Anthropic's own documentation, retrieved 2026-09-08, never from running a workload against any model.
Sources: the Claude API pricing page for the base rate card, subscription tiers, and the managed-agents rate; the platform pricing reference for tool-use overhead and multiplier-stacking rules; the prompt caching guide for per-model cache minimums; and the effort and thinking cost pages for how extended thinking bills. Two figures below are our own arithmetic, labeled as such where they appear. Full methodology: llmwaves.com/methodology.
This piece ranks the four current models on price and billing mechanics, not on output quality.

The September 1 Sonnet 5 price increase never took effect
Claude Sonnet 5 lists at $2 per million input tokens and $10 per million output tokens on the pricing page retrieved 2026-09-08. A separate rate of $3 and $15 was scheduled to begin September 1, 2026. That page does not show it.
The $2 and $10 figures are the ones billed on every Sonnet 5 request today, batch or standard.
Nine days is a short window for a reversed price to fully settle anywhere it was written down. Budgeting from the higher, unrealized rate overshoots a Sonnet 5 forecast by 50%. The fix is one line: check the rate against the model's own page directly, not against a cached screenshot of the announcement.
Tool use adds input tokens before a single call executes
Naming tools in a request adds tokens whether or not the model ever calls one. Anthropic's own reference lists the overhead by model and by tool_choice setting: Claude Opus 5 adds 286 tokens at tool_choice: auto or none, and 406 tokens at any or tool.
Claude Sonnet 5 adds 354 and 474. Claude Haiku 4.5 adds 496 and 588, the largest tax of the three despite being the cheapest per token.
That overhead bills as input tokens, the same bucket as the prompt itself, alongside every tool_use and tool_result block a turn produces.
Server-side tools carry their own separate charges on top: web search runs $10 per 1,000 searches, and code execution runs $0.05 per container-hour after a monthly free allowance.
A request that defines 10 tools and never calls one still pays the definition tax on every single turn.

Extended thinking bills at the output rate, not a premium one
Anthropic's own documentation states thinking tokens are "billed as output tokens," at the same per-token price as the visible answer, not a separate or higher rate. The effort parameter controls how much a turn thinks, with levels from low through max, and the current default is high: setting effort to high explicitly produces the same behavior as leaving the parameter out entirely.
Thinking tokens count toward max_tokens, and the field output_tokens_details.thinking_tokens reports how many of the billed output tokens were spent on internal reasoning rather than the final answer.
There is no thinking surcharge. A bigger budget raises cost only by raising the token count. LLM Waves Research modeled this directly: a fixed 600-token visible answer on Sonnet 5, at output price $10 per million tokens, costs $0.006 with no thinking budget attached.
The same answer with a 4,096-token thinking budget attached costs $0.047, close to 8 times as much for an identical final response. At a 32,768-token budget the same request costs $0.334, roughly 55 times the no-thinking cost.
This ratio is our own calculation, not a published Anthropic figure, and it moves with the visible-answer length used.

Prompt caching only pays off above a per-model token floor
A cache write costs more than a plain input token: 1.25 times the base rate at a 5-minute time to live, 2 times at 1 hour. A cache hit costs far less: 0.1 times the base rate on Haiku 4.5, Sonnet 5, and Opus 5, and 0.025 times on Fable 5.1.
Anthropic's own math says break-even is one cache read for a 5-minute cache and two reads for a 1-hour cache.
None of that math applies below a minimum prompt length, and the minimum is not the same number everywhere. Sonnet 5 needs 1,024 tokens before a prompt is eligible for caching. Opus 5 and Fable 5.1 need only 512. Haiku 4.5 needs 4,096, 4 times the floor set for Sonnet 5 and 8 times the floor set for the two smaller models.
A shorter prompt marked with cache_control is not rejected and does not error.
It is processed at full price with no cache write and no cache read, and the only way to catch this is to check whether cache_creation_input_tokens and cache_read_input_tokens both came back as zero.

A metered API bill crosses a $20 subscription within about 2,200 requests
The Pro subscription runs $17 a month billed annually or $20 a month billed monthly, and it bundles Claude Code rather than selling it separately; Max starts at $100 a month for 5 times Pro's usage, with a heavier tier above it. Neither plan charges per token. The API does, and the two spending models only compare cleanly at a specific volume.
Using the 2,000-input, 500-output token profile carried through this article, one request against Sonnet 5 costs (2,000 times $2 plus 500 times $10) divided by 1,000,000, or $0.009. A $20 monthly spend on the metered API buys roughly 2,200 requests of that shape, about 74 a day.
Below that volume, a subscription seat is the more expensive way to send the same 2,000-and-500-token pattern through Sonnet 5; above it, per-token billing is cheaper unless prompt caching and batching are already stacked into the metered side, which lowers the crossover point further.
Neither plan is the frugal choice at every volume, and the crossover moves with the input-to-output ratio of the actual workload, not with this article's chosen profile.

US-only data residency and platform routing both change the bill
Setting inference_geo to us applies a 1.1 times multiplier to input, output, and both cache states, on Claude 4.6-generation models and newer.
Older models reject the parameter outright. Leaving inference_geo at its global default carries no multiplier at all.
Routing through a different platform changes the calculation entirely rather than adding a fixed multiplier on top of it. Anthropic's own documentation states that Amazon Bedrock and Google Cloud carry independent regional pricing, not the 1.1 times US-residency surcharge described above; a Microsoft Foundry deployment using the US Data Zone Standard option applies that 1.1 times figure the same way the direct API does.
A full rate-by-rate comparison of the direct API against Bedrock and Vertex AI pricing is enough of its own subject to earn a separate article rather than a paragraph here.
What we did not measure
This article did not run a latency test, a throughput test, or an output-quality benchmark against any Claude model. Every number above is a collected rate, a collected mechanic, or an arithmetic derivation clearly marked as our own, never a performance measurement.
Which model for which workload
Developers
Best value: Haiku 4.5 on raw per-token price. Its 4,096-token cache floor and its larger tool-definition overhead both erode that lead on short, tool-heavy prompts.
Startups
Best value: Sonnet 5 with prompt caching and the Batch API stacked together, the combination modeled above at a 65% reduction against full price.
Enterprise
Situational. A data-residency requirement is decided by the 1.1 times US surcharge against whichever regional Bedrock or Vertex rate applies, not by one universal answer.
High volume
Top pick: the Batch API on any of the four models. The 50% discount stacks with prompt caching rather than replacing it.
Low latency
Situational. Fast Mode on Opus 5 runs at roughly 2.5 times normal speed for 2 times the standard token rate, a real lever this piece did not measure against a timer.
Self-hosting
Not applicable. None of the four models ship as downloadable weights.
On-device
Insufficient data. No figures on this page cover on-device inference of a hosted-only model.
Non-English
Insufficient data. Anthropic's pricing pages price by token count regardless of language, and nothing here isolates a non-English cost difference.
Is there a free Claude API tier?
No published free tier exists on claude.com/pricing for ongoing API usage. Anthropic's free-of-charge consumer chat tier applies to the web, iOS, Android, and desktop chat surfaces, not to API keys, and it excludes Claude Code access.
New API accounts sometimes carry promotional credit, which is a temporary balance rather than a standing free tier and is not itemized on the pricing page used for this article.
How does prompt caching actually cut the bill?
It splits a repeated prompt into a write, priced above the standard input rate, and a read, priced at a fraction of it. The worked example above modeled a 1-hour cache write on 1,500 of 2,000 input tokens per request, with 499 of every 500 daily requests landing as cheap reads rather than fresh writes, and that pattern alone cut the monthly bill by 30% before batch was added on top.
What are the minimum cacheable prompt lengths per model?
1,024 tokens for Sonnet 5, 512 for Opus 5 and Fable 5.1, and 4,096 for Haiku 4.5, all stated directly on Anthropic's prompt caching page. A prompt under its model's floor is billed at full price with no error returned.
Does Claude charge a surcharge for US-only data residency?
Yes, on the direct API: 1.1 times the standard rate across input, output, and both cache states, on Claude 4.6-generation models and newer.
That surcharge is specific to Anthropic's own inference_geo parameter and does not carry over to Bedrock or Vertex AI, which price the same models independently by region instead.
Is the Claude API cheaper than OpenAI's or Gemini's?
There is no single number that answers this honestly, because each vendor prices multiple current tiers rather than one flagship rate, and a fair comparison has to match tier to tier rather than picking whichever pair makes the largest gap. A full per-vendor, per-tier breakdown is its own article rather than a paragraph inside this one.
Anthropic will move at least one of these numbers before this quarter ends
Nine days separated the announced Sonnet 5 increase from its cancellation. A model launch, a rate change, or a new effort default can move any figure on this page just as fast. Treat every number above as dated 2026-09-08, not as permanent.
