Skip to content
TokenGauge
Menu

How-to

How to Cut an LLM API Bill: Five Levers Ranked by Impact

Every lever below is priced on the same workload: Claude Sonnet 5 on OpenRouter, 50 million input and 15 million output tokens a month. The baseline is $264.

Updated 9 Oct 2026 · 8 min read · First published 9 Oct 2026

Summary

  • Model choice is the largest lever. Moving the workload to Haiku 4.5 halves the bill.
  • A Haiku-first cascade with 30 percent escalation cuts about 20 percent.
  • Switching gateway to SayGM cuts about 16 percent with no code changes beyond the client.
  • Trimming output by a fifth saves more than trimming input by a fifth.
  • Combined, the levers take the bill from $264 to about $157.

01

The baseline

Claude Sonnet 5 on OpenRouter costs $2.11 per million input tokens and $10.55 per million output tokens, including the top-up fee. At 50M input and 15M output tokens a month, that is $105.50 plus $158.25, or $263.75.

Monthly cost after each lever, applied on its own to the baseline.
LeverMonthly costChange
Baseline$264n/a
Move all traffic to Haiku 4.5$132−50%
Haiku-first cascade, 30% escalate$211−20%
Switch gateway to SayGM$222−16%
Trim output tokens by 20%$232−12%
Trim input tokens by 20%$243−8%

02

1. Choose the smallest model that passes

Haiku 4.5 costs half of Sonnet 5 per token with every provider. Sonnet 5 costs two fifths of Opus 5. No other lever comes close, so test the cheaper model first and keep the larger one for tasks that fail.

Same workload on SayGM.
ModelMonthly costvs Sonnet 5
Claude Opus 5$5562.5x
Claude Sonnet 5$2221.0x
Claude Haiku 4.5$1110.5x

The Sonnet 5 vs Haiku 4.5 comparison shows the gap at three usage levels.

03

2. Try the small model first

A cascade sends every request to Haiku 4.5 and escalates to Sonnet 5 on failure. The worst case is that every escalated request pays for both attempts. Under that assumption the break-even escalation rate is 50 percent, because Haiku costs exactly half.

OpenRouter prices, both attempts billed on escalation.
Escalation rateMonthly costvs all Sonnet
10%$159−40%
30%$211−20%
50%$2640%

SayGM runs cascades at the gateway level, so no routing code is needed. On OpenRouter you write the retry yourself.

04

3. Stop paying for the gateway

The same Sonnet 5 workload costs $222.35 on SayGM: $89 for input and $133.35 for output. That is $41.40 a month, or $497 a year, for a base URL change.

Fig. 1Monthly bill at three usage levels, Claude Sonnet 5
  • OpenRouter, fee included
  • SayGM
Side project5M in · 1M out
$21
$18
Production app50M in · 15M out
$264
$222
High volume500M in · 100M out
$2,110
$1,779
Show data as a table
OpenRouter, fee includedSayGM
Side project (5M in · 1M out)$21$18
Production app (50M in · 15M out)$264$222
High volume (500M in · 100M out)$2,110$1,779
Usage is millions of input and output tokens per month. No caching discount applied. Source: published rate cards, October 2026 snapshot. Current prices on saygm.com (opens in a new tab).

The gap is wider on other vendors: about 29 percent on GPT-5.5 and 34 percent on Gemini 3.1 Pro. See all prices.

05

4. Trim output before input

Output tokens cost five times more than input on Claude. At an 80/20 mix, the 20 percent of tokens a model writes are 56 percent of the bill.

Fig. 2Share of the bill from output tokens at an 80/20 token mix
  • Input tokens (80% of volume)
  • Output tokens (20% of volume)
Claude Sonnet 5
Claude Haiku 4.5
GPT-5.5
Gemini 3.1 Pro Preview
Show data as a table
ModelShare of bill from inputShare of bill from output
Claude Sonnet 544%56%
Claude Haiku 4.544%56%
GPT-5.540%60%
Gemini 3.1 Pro Preview40%60%
Output tokens are a fifth of the volume but most of the cost on every frontier model. Trimming output length saves more than trimming prompts. Source: published rate cards, October 2026 snapshot. Current prices on saygm.com (opens in a new tab).

Cutting output by 20 percent removes 3M tokens and saves $31.65 a month. Cutting input by 20 percent removes 10M tokens and saves $21.10. Ways to cut output:

  • Set max_tokens per task, not one global ceiling.
  • Ask for diffs or edits instead of full rewritten files.
  • Return structured output with short field names.
  • Drop restated context and closing summaries from system prompts that ask for them.

06

5. Keep the cache warm

Prompt caching discounts repeated prefixes such as system prompts, tool definitions and long documents. The saving depends on the provider's cache pricing and your hit rate, so it is not in the table above.

  • Put stable content first and variable content last.
  • Keep tool definitions in a fixed order.
  • Keep a conversation on one provider. Gateways that spread a conversation across providers lose cache hits. SayGM pins a conversation to one provider.

07

All levers together

Apply three levers at once: SayGM as the gateway, output trimmed to 12M tokens, and a Haiku-first cascade with 30 percent escalation. Haiku costs $97.90 for the full volume. Escalations add 30 percent of the trimmed Sonnet cost of $195.68, or $58.70.

Run the numbers for your own volumes in the calculator. Rates are a snapshot; SayGM's pricing page (opens in a new tab) has the current figures.

Prices in this guide are a October 2026 snapshot and move each epoch.

Current SayGM pricing (opens in a new tab)