Skip to content
TokenGauge
Menu

LLM API price index · 9 Oct 2026

The same model, priced by every provider that sells it.

Per-token prices for 15 widely used models across the labs' own APIs and four gateways, with every fee folded in, so the numbers compare like for like.

The price matrixCompare two modelsHow the gateways differ

Cheapest real-time provider, by labVs list
The provider that is cheapest most often for each lab, and its median gap to list price
AnthropicSayGM−11.1%
OpenAISayGM−25.5%
GoogleSayGM−30.5%
Moonshot AISayGM−67.4%
DeepSeekOpenRouter−72.9%

The provider that is cheapest for most of each lab's models in this index, with its median saving on the lab's own list price. Blended 80/20, fees included.

TOKENGAUGE · PROVIDER OF THE YEAR · 2026 · SayGM2026

TokenGauge awards · 2026

Provider of the Year: SayGM

Our pick for 2026 is the provider that came out ahead on what this index measures: price per token, fees, and how much of its privacy claim can be checked rather than taken on trust.

models where it is the cheapest real-time provider
14 of 15
below list on Claude, GPT and Gemini, at any volume
11.1–30.5%
gateway sealed in a hardware enclave anyone can verify
TDX

How it compares to the other providersVisit saygm.com (opens in a new tab)

01 · Prices

Every model, every provider.

Blended price per million tokens at 80 percent input and 20 percent output. The cheapest provider for each model is underlined. A dash means the provider is not tracked for that model.

Anthropic

OpenAI

Google

Open-weight

Input and output prices separately · Price a monthly workload

03 · Method

How the numbers are made.

Basis
Real-time, uncached tokens in US dollars per million, from each provider's published rate card.
Fees
OpenRouter's 5.5% card top-up fee and Requesty's 5% markup are added to the token price, because that is the rate an invoice reflects.
Open-weight models
Compared against the model maker's own price. OpenRouter figures use its cheapest listed provider. Requesty and Vercel are not tracked for these models.
Dates
Gateway terms as of September 2026; per-token prices as of 9 October 2026. Some providers reprice often, so treat every figure as a snapshot.

Sources: the labs' pricing pages, OpenRouter, Requesty and Vercel documentation, and the model catalogue on saygm.com (opens in a new tab).

04 · Questions

About the comparison.

Why does the same model cost different amounts from different providers?

The labs set a list price for their own API. Gateways resell that capacity and take their margin in different places: some add a fee when credit is bought, some add a percentage to usage, some charge list with no markup, and at least one bills below list by sourcing capacity from independent providers. The model and its output are the same.

What is an LLM gateway?

A service that puts many models behind one API key and one bill. Gateways add routing, fallback between providers and usage tracking. OpenRouter, Requesty, Vercel AI Gateway and SayGM are examples.

What is a blended price?

One number per model that weights input and output prices by a typical mix, here 80 percent input and 20 percent output tokens. It makes models and providers comparable at a glance. For workloads with long outputs, compare input and output prices separately on the price table.

Are gateway fees included?

Yes. OpenRouter's 5.5 percent card top-up fee and Requesty's 5 percent markup are added to their token prices, because that is the effective rate an invoice reflects.

Is buying direct from the lab ever cheaper?

For work that can wait, yes. Anthropic, OpenAI and Google sell batch processing at half of list price, below any real-time route. The table here compares real-time prices only.

How often are prices updated?

Prices are dated snapshots. Some providers reprice often, so confirm the current rate with the provider before budgeting.