Across 15 current models from six vendors, output prices run from $0.50 per 1M tokens on GPT-6 Luna to $50 on Claude Fable 5.1 and GPT-6 Astra, a 100x spread.

The cheapest flagship-class output is DeepSeek-V4-Pro at $1.98 off-peak. Prices collected October 1, 2026. The table below holds the full answer.

Current API pricing, USD per 1M tokens

Model

Vendor

Input

Output

Context window (tokens)

Claude Fable 5.1

Anthropic

$10.00

$50.00

1,000,000

GPT-6 Astra

OpenAI

$10.00 (†$20.00)

$50.00 (†$75.00)

1,050,000

Claude Opus 5.5

Anthropic

$4.00

$20.00

1,000,000

Gemini 3.1 Pro Preview

Google

$2.00 (¶$4.00)

$12.00 (¶$18.00)

1,048,576

GPT-6.1 Sol

OpenAI

$2.00 (†$4.00)

$10.00 (†$15.00)

1,050,000

Claude Sonnet 5.5

Anthropic

$2.00

$10.00

1,000,000

Qwen3.8-Max

Alibaba

$2.00

$6.00

1,000,000

Claude Haiku 4.5

Anthropic

$1.00

$5.00

200,000

DeepSeek-V4-Pro

DeepSeek

$1.32 (§$0.66)

$3.96 (§$1.98)

1,000,000

Gemini 3.8 Flash

Google

$0.75 (‡)

$3.75 (‡)

1,048,576

Gemini 3.1 Flash-Lite

Google

$0.25

$1.50

1,048,576

Mistral Large 3

Mistral

$0.50

$1.50

256,000

DeepSeek-V4.1-Flash

DeepSeek

$0.30 (§$0.15)

$1.20 (§$0.60)

1,000,000

Mistral Small 4

Mistral

$0.15

$0.60

256,000

GPT-6 Luna

OpenAI

$0.10 (†$0.20)

$0.50 (†$0.75)

1,050,000

Table note: rows run by output price, highest first. † OpenAI long-context rate, charged on the whole request once input passes the long-context threshold. ¶ Google rate above 200,000 input tokens. § DeepSeek off-peak rate; the main figure is the peak rate. ‡ Google introductory rate through 2026-12-31, rising to $1.50 input and $7.50 output on 2027-01-01. All figures are vendor list prices for direct API access, collected on 2026-10-01 UTC. None were produced by sending requests to these APIs.

Claude Fable 5.1 and GPT-6 Astra share the top output rate, $50. GPT-6 Luna sits at the bottom, at $0.50.

That is a 100x spread, and OpenAI sells both ends of it.

Where these numbers come from

Every price above comes from the vendors' own pricing and model documentation, collected on October 1, 2026, and LLM Waves Research did not send requests to any of these APIs. The sources are the Anthropic Claude Platform pricing page, the OpenAI developer pricing page, the Google Gemini API pricing page, the DeepSeek pricing page, the Alibaba Cloud Model Studio pricing page, and the Mistral API pricing page, with the full dated list on the methodology page. LLM Waves Research received no pre-release access, free credits or rate-limit exemptions from any vendor, and no vendor saw this comparison before publication.

The ranking runs on output price per 1M tokens first, then on context window, cache discount, and free allocation.

How LLM API pricing actually works

An API request carries two prices. Input tokens are the prompt, the system instructions, and any retrieved context.

Output tokens are everything the model generates, billed at three to six times the input rate across this table.

A 1,000-token question that returns a 2,000-token answer from Claude Sonnet 5.5 costs $0.002 for input and $0.020 for output. Output is ten times the input side on that one call.

Caching discounts the input side only. A cache hit reuses tokens already processed, usually a repeated system prompt or a long document, at a fraction of the normal input rate. The discount is no longer a single ratio:

  • DeepSeek-V4.1-Flash: $0.003 cache hit against $0.15 off-peak, a 50x cut

  • Claude Fable 5.1: $0.25 against $10.00, a 40x cut

  • DeepSeek-V4-Pro: $0.022 against $0.66 off-peak, a 30x cut

  • Claude Opus 5.5 and GPT-6.1 Sol: $0.20 against $4.00 and $0.10 against $2.00, both 20x

  • Claude Sonnet 5.5, GPT-6 Astra and GPT-6 Luna: 10x

Anthropic prices most Claude cache hits at 0.1x base input. Claude Fable 5.1 and Claude Opus 5.5 are the exceptions, which is why the two most expensive Claude tiers carry the deepest Claude cache discounts.

Batching trades speed for price. The Claude Opus 5.5 batch rate is $2 input and $10 output, half the standard rate, and Google applies the same 50% batch discount across the Gemini lineup.

Neither caching nor batching changes what a model produces. Both need a workload built for them.

Reasoning tokens are output tokens the reader never sees

OpenAI reasoning models generate an internal reasoning pass before the visible answer, billed at the output rate even though the requester never reads it.

The o-series shows the effect in its own price list: o3 lists $8 output, o3-pro lists $80, and o1 lists $60 against $10 for GPT-6.1 Sol.

Two requests with the same prompt length can differ tenfold in cost once the reasoning setting moves up.

The setting decides how many hidden output tokens run before the first visible word.

Context window is a separate purchase from price

A low output rate says nothing about how much text a model accepts. Mistral Large 3 and Mistral Small 4 both stop at 256,000 tokens, and Claude Haiku 4.5 stops at 200,000, against roughly 1,000,000 for every other model in the table.

All three OpenAI tiers share the same 1,050,000-token window, from GPT-6 Astra at $50 output down to GPT-6 Luna at $0.50. Window size does not track price inside the OpenAI lineup at all.

The three Gemini models share a 1,048,576-token input limit and a 65,536-token output limit, half the 128,000-token output ceiling listed for Claude Fable 5.1, Claude Opus 5.5 and Claude Sonnet 5.5.

Price thresholds sit inside some windows. Gemini 3.1 Pro Preview doubles input to $4.00 and raises output to $18.00 above 200,000 input tokens, and the OpenAI long-context rate applies to the whole request once it crosses the threshold.

A workload built around long documents can rule out a cheap model on window size before price enters the decision.

What is the cheapest LLM API, flagship tier, and budget tier

OpenAI lists three current tiers on its developer pricing page: GPT-6 Astra at $10 and $50, GPT-6.1 Sol at $2 and $10, and GPT-6 Luna at $0.10 and $0.50.

GPT-5.6 Sol stays on sale at $4 and $20 under promotional pricing that OpenAI says runs at least through November 21, 2026.

Anthropic leads with Claude Fable 5.1 at $10 and $50, then Claude Opus 5.5 at $4 and $20 and Claude Sonnet 5.5 at $2 and $10.

Claude Opus 5 ($5 and $25) and Claude Sonnet 5 ($2 and $10) remain on sale under the additional models list.

DeepSeek prices by time of day. The DeepSeek pricing page sets peak hours at 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, and charges half the peak rate at every other hour.

DeepSeek-V4-Pro output costs $3.96 at peak and $1.98 off-peak. Weekends are off-peak in full.

At the flagship tier, DeepSeek-V4-Pro off-peak is the floor: $1.98 output against $50 for Claude Fable 5.1 and GPT-6 Astra, a 25x gap. At peak the gap narrows to 12.6x.

The budget tier reverses the usual story. GPT-6 Luna lists $0.50 output, below DeepSeek-V4.1-Flash at $0.60 off-peak and $1.20 at peak, Mistral Small 4 at $0.60, and Gemini 3.1 Flash-Lite at $1.50.

Claude Haiku 4.5 sits highest of the budget tiers at $5.00. The full DeepSeek peak schedule is broken down in the DeepSeek API pricing guide.

Which LLM gives the best price-to-performance

Price alone cannot say which model is worth paying for, and this article runs no quality benchmark of its own. The pricing data shows the shape of the trade: GPT-6.1 Sol and Claude Sonnet 5.5 both list $2 input and $10 output, Gemini 3.1 Pro Preview lists $2 and $12, and Qwen3.8-Max lists $2 and $6. Four models from four vendors sit at the same input rate.

A reader who needs quality scores next to these prices should treat that as a second lookup.

The model index carries those scores; this page carries the price, half of the decision.

What are the real rate limits behind the price per token

A price per token says nothing about whether an account can run production volume, and none of the three largest vendors publish per-tier request ceilings before signup.

Anthropic names three usage tiers, Start, Build and Scale, plus a Custom tier for managed accounts, measured in requests, input tokens and output tokens per minute, according to its rate-limits article.

Higher limits can be requested once an account uses at least 50% of its current ones, and the numbers sit inside the Claude Console.

OpenAI points readers to the limits page inside account settings, where an organization applies for an increase, per its help center article on rate limits.

Google publishes the most. The Gemini rate-limits documentation sets spend caps of $10, $50, and $200 per 10 minutes at Tier 1, Tier 2, and Tier 3, with Tier 2 opening after $100 paid plus 3 days and Tier 3 after $1,000 paid plus 30 days.

Request and token ceilings per tier appear only in Google AI Studio.

A $0.50 output price and a $50 output price share the same units, but neither says whether the account behind it can send ten requests a second or one.

Which providers offer a free tier, and how much

Three of the six vendors list something free on the API itself.

The Google Gemini API pricing page offers a free tier on Gemini 3.8 Flash and Gemini 3.1 Flash-Lite and none on Gemini 3.1 Pro Preview.

Alibaba Cloud gives each model an independent free quota: 1,000,000 tokens in the examples its free-quota documentation uses, valid for 90 days from activation or model release, whichever is later. Qwen3.8-Max is covered.

Mistral lists two free endpoints on its API pricing page: Mistral Moderation 2 and Leanstral 1.5. Mistral Large 3, Mistral Medium 3.5 and Mistral Small 4 are paid from the first token.

DeepSeek lists no free allocation on its pricing page, though its deduction rules mention a granted balance spent before a topped-up one.

OpenAI and Anthropic publish pay-as-you-go rates from the first token.

How do you know a pricing page is current?

Two checks take under a minute, both against the vendor site rather than the page making the claim.

First, match the model ID, not the marketing name. The current IDs in this table are gpt-6-astra, gpt-6.1-sol, gpt-6-luna, claude-fable-5-1, claude-opus-5-5, claude-sonnet-5-5, gemini-3.1-pro-preview, gemini-3.8-flash, deepseek-v4-pro, deepseek-flash, qwen3.8-max and mistral-large-2512. The DeepSeek page notes that the old names deepseek-v4-flash and deepseek-v4-flash-vision-exp still work, but those models are retired, and requests run on DeepSeek-V4.1-Flash at the Flash price.

Second, look for an end date inside the price. Gemini 3.8 Flash doubles to $1.50 and $7.50 on January 1, 2027. GPT-5.6 Sol carries a promotional rate through at least November 21, 2026.

Six of the 15 rows in this table changed between September 7 and October 1, 2026. GPT-6.1 Sol and GPT-6 Luna replaced GPT-5.6 Sol and GPT-5.6 Luna at lower rates, Claude Opus 5.5 replaced Claude Opus 5 at a lower rate, Claude Sonnet 5.5 replaced Claude Sonnet 5 at the same rate, and DeepSeek-V4.1-Flash replaced DeepSeek-V4-Flash.

The DeepSeek-V4-Pro row moved to the peak and off-peak rates DeepSeek lists, $3.96 and $1.98 output.

When does self-hosting beat paying per token

Self-hosting replaces the per-token bill with a GPU rental or ownership cost that does not scale down with traffic.

Mistral Large 3 and Mistral Small 4 carry Apache 2.0 licences on the Mistral pricing page, so both can run on rented or owned hardware.

Below some request volume, a rented GPU running around the clock costs more than any per-token rate in the table.

Above it, hardware cost stays flat while the token count climbs. Where that crossover sits depends on utilization, and estimating it needs an hourly hardware rate this article does not collect.

What we did not measure

This article did not run a latency test, a throughput test, or an output-quality benchmark against any model discussed. It reports public list prices only: no negotiated enterprise rates, no committed-use discounts, no third-party inference marketplace prices, and no tokenizer comparison across languages.

It reports no numeric rate-limit ceiling for Anthropic or OpenAI, since neither publishes one outside the account console.

Use-case verdicts

If you are

Pick

Price per 1M tokens (input/output)

Why

When it is the wrong pick

A developer testing one model

GPT-6 Luna (best value)

$0.10 / $0.50

Lowest price pair in the table, with a 1,050,000-token window

You need an open-weight model. Look at Mistral Small 4

A startup that can run jobs at any time

DeepSeek-V4.1-Flash (situational)

$0.15 / $0.60 off-peak

Cheap if your jobs run outside DeepSeek's busy hours

Your jobs run in busy (peak) hours, when both prices double

A big company at the mid tier

GPT-6.1 Sol or Claude Sonnet 5.5 (tie)

$2 / $10 for both

Same price, so neither is ahead

GPT-6.1 Sol reads 1,050,000 tokens; Claude Sonnet 5.5 reads 1,000,000

Sending a high volume of flagship-class requests

DeepSeek-V4-Pro (top pick on price)

$1.98 output off-peak

Every other flagship-class model costs $6 to $50 output

Peak traffic doubles the output price to $3.96

Needing fast replies (low latency)

Not enough data

None

None of the six pricing pages lists a latency figure

None

Running on your own device

Not enough data

None

Every model here is sold as a hosted API

None

Self-hosting

Mistral Large 3 or Mistral Small 4 (situational)

$1.50 and $0.60 output

Apache 2.0 licences, so you can run them on your own hardware

Your volume is too low for hardware to cost less than $1.50 and $0.60 output

Working in languages other than English

Not enough data

None

No pricing page here lists tokens per word by language

None

How much does 1 million tokens cost?

It depends on the side of the request and the model. One million output tokens cost $0.50 on GPT-6 Luna and $50 on Claude Fable 5.1 or GPT-6 Astra.

Input costs three to six times less than output at every vendor in the table.

Is there a free LLM API?

Not a standing one from all six vendors, but three come close. Google offers a free tier on Gemini 3.8 Flash and Gemini 3.1 Flash-Lite.

Alibaba Cloud gives each model a 90-day free quota. Mistral keeps Mistral Moderation 2 and Leanstral 1.5 free.

How does cached-input pricing change the real cost?

Cached input costs 10x to 50x less than standard input in this table, while output stays at full price. DeepSeek-V4.1-Flash has the deepest cut at 50x.

A workload that repeats a long system prompt gains far more than one built from unrelated prompts.

Is DeepSeek's API cheaper than OpenAI's?

At the flagship tier, yes. DeepSeek-V4-Pro lists $1.98 off-peak output against $50 for GPT-6 Astra and $10 for GPT-6.1 Sol.

At the budget tier, no: GPT-6 Luna lists $0.50 output, under DeepSeek-V4.1-Flash at $0.60 off-peak and $1.20 at peak.

How do reasoning tokens affect cost?

Reasoning tokens bill as output tokens even though the requester never sees them.

OpenAI prices o3-mini at $4.40 output, o3 at $8 and o3-pro at $80, an 18x range inside one model family.

Three prices in this table carry an end date

Gemini 3.8 Flash doubles on January 1, 2027. GPT-5.6 Sol promotional pricing runs at least through November 21, 2026. DeepSeek charges double during 7 weekday hours.

Before committing a budget to any row, re-check the vendor page named in the methodology; the changelog below records each revisit of this table.