Skip to content
TokenGauge
Menu

DeepSeek · Open weight · TEE variant available

DeepSeek V4.1 Flash API pricing

40 percent below DeepSeek's list price on SayGM's open tier. OpenRouter's cheapest provider is cheaper still, but not confidential. The sealed -tee variant runs inside an Intel TDX enclave.

Snapshot October 2026 · USD per million tokens · Best for: bulk generation, cheap open-weight inference
SayGM, in / out
$0.18 / $0.72
No top-up fee
OpenRouter, in / out
$0.03 / $0.53
Cheapest OpenRouter provider including the 5.5 percent fee
Output to input price
4.0x
On SayGM
Per 100M tokens, 80/20
$29
$13 on OpenRouter

§ 01

What a month costs

The same three workloads used across the index. Here the cheapest OpenRouter provider comes out slightly lower at an 80/20 mix.

Fig. 1Monthly bill at three usage levels, DeepSeek V4.1 Flash
  • OpenRouter, fee included
  • SayGM
Side project5M in · 1M out
$0.68
$1.62
Production app50M in · 15M out
$9.45
$20
High volume500M in · 100M out
$68
$162
Show data as a table
OpenRouter, fee includedSayGM
Side project (5M in · 1M out)$0.68$1.62
Production app (50M in · 15M out)$9.45$20
High volume (500M in · 100M out)$68$162
Usage is millions of input and output tokens per month. No caching discount applied. Source: published rate cards, October 2026 snapshot. Current prices on saygm.com (opens in a new tab).

§ 02

Where the money goes

Output tokens cost 4.0 times more than input on DeepSeek V4.1 Flash. At an 80/20 mix they make up 50% of the bill.

Fig. 2Share of the bill from output tokens at an 80/20 token mix
  • Input tokens (80% of volume)
  • Output tokens (20% of volume)
DeepSeek V4.1 Flash
Kimi K3
Show data as a table
ModelShare of bill from inputShare of bill from output
DeepSeek V4.1 Flash50%50%
Kimi K345%55%
Output tokens are a fifth of the volume but most of the cost on every frontier model. Trimming output length saves more than trimming prompts. Source: published rate cards, October 2026 snapshot. Current prices on saygm.com (opens in a new tab).

§ 03

Estimate a DeepSeek V4.1 Flash bill

50M
10M

Monthly bill by provider

  • OpenRouter, fee included$7
  • DeepSeek direct, list price$27
  • SayGM$16
A year, vs the cheapest alternative
+$113
138% higher than OpenRouter
A year, vs OpenRouter
+$113
OpenRouter's cheapest provider costs less

Uncached prices from the October 2026 snapshot. Prompt caching lowers every provider's bill; batch jobs bought direct cost half of list. Requesty and Vercel are not priced for open-weight models here. Current SayGM rates are on saygm.com (opens in a new tab).

§ 04

Other open-weight models

USD per million tokens, input and output. OpenRouter includes its 5.5% top-up fee. Snapshot from October 2026.
ModelSayGMin / outOpenRouterin / out, fee incl.SayGM vs OpenRouter80/20 blend
Kimi K3Moonshot AI · TEE variant$0.98 / $4.88$0.42 / $9.50Cheapest provider
21% lower
DeepSeek V4.1 FlashDeepSeek · TEE variant$0.18 / $0.72$0.03 / $0.53Cheapest provider
122% higher

§ 05

Compare and set up

§ 06

DeepSeek V4.1 Flash questions

How much does DeepSeek V4.1 Flash cost on SayGM?

As of October 2026, $0.18 per million input tokens and $0.72 per million output tokens, with no top-up fee. Prices are set per epoch and capped at list, so check saygm.com for the live rate.

Is DeepSeek V4.1 Flash cheaper on SayGM or OpenRouter?

OpenRouter's cheapest provider is slightly cheaper at an 80/20 mix for this model, mainly on input tokens, but it is not served inside a trusted execution environment. SayGM is cheaper on output tokens.

Which API does SayGM use for DeepSeek V4.1 Flash?

The model's native API. DeepSeek features such as caching, tool use and structured output work unchanged. Any OpenAI-compatible SDK works by changing the base URL.