Claude Opus 5 lists at $5 per million input tokens and $25 per million output tokens, a flat rate across its full 1,000,000-token window. Gemini 3.1 Pro Preview starts cheaper, at $2 and $12. That rate doubles to $4 and $18 once a prompt crosses 200,000 tokens.

Short prompts favor Gemini 3.1 Pro by 2.5x on input and roughly 2.1x on output. Long prompts favor it by only 1.25x and 1.4x.

Gemini 3.1 Pro also carries a Preview label, with no production SLA, as of September 11, 2026. Gemini 2.5 Pro is the only Pro-tier Gemini model Google lists as generally available.

Model

Input, USD/1M

Output, USD/1M

Context window

Max output

Status

Claude Opus 5

$5.00 (flat)

$25.00 (flat)

1,000,000 tokens

128,000 tokens

GA

Claude Sonnet 5

$2.00 (flat)

$10.00 (flat)

1,000,000 tokens

128,000 tokens

GA

Claude Haiku 4.5

$1.00 (flat)

$5.00 (flat)

200,000 tokens

64,000 tokens

GA

Gemini 3.1 Pro

$2.00 / $4.00*

$12.00 / $18.00*

1,000,000 tokens

65,536 tokens

Preview

Gemini 2.5 Pro

$1.25 / $2.50*

$10.00 / $15.00*

1,000,000 tokens

unconfirmed†

GA

* Lower rate applies at or below 200,000 prompt tokens; higher rate applies above it. † Not published in the sources checked for this article; excluded from ranking rather than guessed.

Basis and method

Every figure in this article was collected, not tested. LLM Waves Research ran no benchmark and executed no API call against either model for this comparison.

Pricing for Claude was collected from Anthropic's current rate card on September 11, 2026. Benchmark and status figures came from each vendor's own launch materials and model cards, supplemented by one named third-party tracker where a vendor's own materials published no comparable number.

Pricing for Gemini was collected from Google's Gemini API pricing documentation on the same date. Arithmetic performed by LLM Waves Research is limited to the price-ratio calculations shown here: the higher-priced model's rate divided by the lower-priced model's rate, at matching input or output position. This article is Measurement Basis 2, aggregated. No figure here should be read as a "we tested" result. Full source list: methodology.

The comparison ranks Claude Opus 5 against Gemini 3.1 Pro Preview as the current flagship tier from each vendor, with Claude Sonnet 5, Claude Haiku 4.5, and Gemini 2.5 Pro included in the use-case verdicts below. Full specifications for the Claude side sit on the Claude Opus 5 model page.

Gemini 3.1 Pro's short-context rate beats Claude Opus 5 by more than double on both input and output. Cross 200,000 prompt tokens and that lead collapses to 25% more on input, 40% more on output.

The threshold nobody prices around

Long-running agent sessions push past 200,000 tokens routinely. So do large codebases pasted into one prompt, and multi-document research tasks. Once a call crosses that line, Gemini 3.1 Pro's own rate card stops being the bargain its headline price suggests. The discount ends at 200,000 tokens.

A 300,000-token research prompt bills at $4 per million input tokens and $18 per million output tokens on Gemini 3.1 Pro, against Claude Opus 5's unchanged $5 and $25. The dollar difference on one such call is small. Repeated hundreds of times a day inside an agent loop, it is not.

Claude Opus 5 carries no long-context pricing tier at all, a fact stated plainly on Anthropic's own pricing page. The rate at token one and the rate at token 999,999 are identical. That design trades away the short-prompt discount Gemini offers, in exchange for a flat, predictable ceiling.

Search and grounding cost more than the tokens do

Claude's API exposes a web search tool at $10 per 1,000 searches, with no free allowance at any volume. Google's equivalent, Search grounding, is free up to a threshold, according to Google's own pricing documentation: 1,500 requests per day on Gemini 2.5 Pro, or 5,000 requests per month shared across all Gemini 3.x models.

Past that threshold, the two Gemini tiers charge different rates for the same feature. Gemini 2.5 Pro bills $35 per 1,000 grounded requests once its daily allowance runs out. Gemini 3.1 Pro bills $14 per 1,000 once its monthly allowance runs out.

A low-volume integration pays nothing on Gemini and $10 flat on Claude. A high-volume one on Gemini 2.5 Pro can end up paying more per request than Claude charges from the first search onward. Price the tool call, not only the token.

Gemini 2.5 Pro's post-free-tier grounding rate runs three and a half times Claude's flat web search price. Gemini 3.1 Pro's rate sits below Claude's, once its larger free allowance runs out.

Coding: the one shared benchmark isn't a controlled test

Anthropic's own launch material for Claude Opus 5 names Frontier-Bench v0.1, CursorBench 3.2, ARC-AGI-3, Zapier AutomationBench, OSWorld 2.0, GDPval-AA v2, and several smaller evaluations. It does not report a SWE-bench Verified score anywhere in that announcement.

Google's launch material for Gemini 3.1 Pro leads with a different benchmark generation entirely: ARC-AGI-2, where it lists 77.1%, not the ARC-AGI-3 suite Anthropic uses. The two companies are not running the same test. Not even the same benchmark generation.

Google's own model card, published separately from the launch announcement, does list a SWE-bench Verified score for Gemini 3.1 Pro: 80.6%, alongside 54.2% on the public SWE-bench Pro split.

Anthropic has not published a matching figure for Claude Opus 5 on either benchmark. A third-party tracker, BenchLM.ai, lists Claude Opus 5 at 96.0% on SWE-bench Verified, a figure Anthropic has neither confirmed nor disputed.

Treating that gap as a verified 15.4-percentage-point win for Claude would overstate what two mismatched sourcing tiers can support.

The closest thing to a shared benchmark between the two vendors' own primary materials is BrowseComp, an agentic search evaluation both companies report separately in their model card and system card. See the Grok vs Claude comparison for the same non-overlap problem against a third vendor.

Gemini 3.1 Pro's 80.6% comes from Google's own model card. Claude Opus 5's 96.0% comes from a tracker neither company runs. The 15.4-percentage-point gap between them mixes two different kinds of evidence.

The output ceiling runs the opposite way from the input one

Claude Opus 5 and Claude Sonnet 5 both generate up to 128,000 tokens in a single response, per Anthropic's context-window documentation. Gemini 3.1 Pro's documented output ceiling is 65,536 tokens, roughly half of Claude's, despite Gemini's input context window matching Claude's at 1,000,000 tokens.

Claude Haiku 4.5 caps output at 64,000 tokens. That puts Claude's cheapest, smallest model close to Gemini 3.1 Pro's ceiling, despite sitting two tiers below it in price.

A task that needs one long response back, a full document rewrite or a large generated file returned in one call, hits Gemini 3.1 Pro's ceiling before it hits Claude Opus 5's.

Claude Opus 5 and Claude Sonnet 5 both generate twice the output of Gemini 3.1 Pro in one response, despite matching input context windows.

Where the models run

Claude has both a global endpoint and general-availability multi-region endpoints on Vertex AI, confirmed on Google Cloud's own blog. A region-pinned option matters for data-residency rules, and it exists here for a model built by a company other than Google.

Whether Gemini 3.1 Pro carries the same regional flexibility on Google's own platform was not confirmed to the same standard this pass. Specifications for that model sit on the Gemini 3.1 Pro model page.

What we did not measure

This article did not run a latency test, a throughput test, or an output-quality benchmark against any model discussed. Every score reported here comes from a vendor's own published material or from a named third-party tracker, never from a request LLM Waves Research sent to either API.

Which model for which reader

Developers building agents

Situational. Below 200,000 tokens per call, Gemini 3.1 Pro's input rate wins outright. Above it, run the arithmetic for your own prompt length before assuming that holds.

Startups watching burn rate

Best value: Gemini 3.1 Pro, for workloads under the 200,000-token threshold that don't lean on the search tool at high volume.

Enterprise deployment

Situational. Claude Opus 5 carries no Preview label and no long-context pricing cliff. Gemini 3.1 Pro carries both, a real consideration for production rather than prototyping.

High-volume production traffic

Situational. Grounding and web-search fees can outweigh token costs at scale on either side. Price the tool calls before the tokens.

Low-latency applications

Insufficient data. Neither vendor's documentation checked here publishes a directly comparable time-to-first-token figure.

On-device or local deployment

Not recommended, either model. Neither Claude Opus 5 nor Gemini 3.1 Pro ships open weights.

Self-hosting

Not recommended, either model. Both are closed-weight and API-only, so neither ships anything to self-host.

Non-English use

Insufficient data. No comparable non-English accuracy figure was published by either vendor in the material checked here.

Is Gemini actually cheaper than Claude?

Below 200,000 prompt tokens, yes. Gemini 3.1 Pro's $2 and $12 rates beat Claude Opus 5's flat $5 and $25 by a wide margin. Above that threshold, Gemini's own rate card doubles, narrowing the gap to 1.25 times on input and 1.4 times on output.

Neither vendor's tool-call pricing follows the same pattern. Check the grounding and search-fee section before assuming the token price is the whole bill.

Which is better for coding, Gemini or Claude?

Anthropic and Google do not publish a shared, directly comparable coding benchmark for these two models. Google's own model card reports 80.6% for Gemini 3.1 Pro on SWE-bench Verified. A third-party tracker, not either vendor, reports 96.0% for Claude Opus 5 on the same benchmark. Those two numbers come from different sourcing tiers. Neither should be read as a controlled head-to-head result.

Is Gemini 3.1 Pro safe to use in production?

Gemini 3.1 Pro carries a Preview label as of September 11, 2026, with no production SLA published alongside it. Gemini 2.5 Pro is the Gemini Pro-tier model Google lists as generally available.

A team that needs a committed support and stability posture from a Pro-tier Gemini model has one GA option, and it is not Gemini 3.1 Pro.

Can Claude and Gemini benchmark scores be compared directly?

Not on the evidence either company has published for its current flagship. Claude Opus 5's launch material and Gemini 3.1 Pro's launch material each headline a different benchmark suite, including two different generations of ARC-AGI.

The one evaluation both vendors' own primary materials report, BrowseComp, is an agentic search test, not a coding or reasoning one. Neither company published it in a format directly comparable to the other's in the materials checked here.

This comparison holds until either company's next release, which could be any week

Claude Opus 5 shipped July 24, 2026. Gemini 3.1 Pro's model card is dated February 19, 2026, and the model was still labeled Preview as of this article's collection date.

Both companies have shipped a new flagship inside three months of a prior one, more than once this year. Recheck the pricing pages linked throughout this article before building a cost model on the numbers above.