Summary
- Model choice is the largest lever. Moving the workload to Haiku 4.5 halves the bill.
- A Haiku-first cascade with 30 percent escalation cuts about 20 percent.
- Switching gateway to SayGM cuts about 16 percent with no code changes beyond the client.
- Trimming output by a fifth saves more than trimming input by a fifth.
- Combined, the levers take the bill from $264 to about $157.
01
The baseline
Claude Sonnet 5 on OpenRouter costs $2.11 per million input tokens and $10.55 per million output tokens, including the top-up fee. At 50M input and 15M output tokens a month, that is $105.50 plus $158.25, or $263.75.
| Lever | Monthly cost | Change |
|---|---|---|
| Baseline | $264 | n/a |
| Move all traffic to Haiku 4.5 | $132 | −50% |
| Haiku-first cascade, 30% escalate | $211 | −20% |
| Switch gateway to SayGM | $222 | −16% |
| Trim output tokens by 20% | $232 | −12% |
| Trim input tokens by 20% | $243 | −8% |
02
1. Choose the smallest model that passes
Haiku 4.5 costs half of Sonnet 5 per token with every provider. Sonnet 5 costs two fifths of Opus 5. No other lever comes close, so test the cheaper model first and keep the larger one for tasks that fail.
| Model | Monthly cost | vs Sonnet 5 |
|---|---|---|
| Claude Opus 5 | $556 | 2.5x |
| Claude Sonnet 5 | $222 | 1.0x |
| Claude Haiku 4.5 | $111 | 0.5x |
The Sonnet 5 vs Haiku 4.5 comparison shows the gap at three usage levels.
03
2. Try the small model first
A cascade sends every request to Haiku 4.5 and escalates to Sonnet 5 on failure. The worst case is that every escalated request pays for both attempts. Under that assumption the break-even escalation rate is 50 percent, because Haiku costs exactly half.
| Escalation rate | Monthly cost | vs all Sonnet |
|---|---|---|
| 10% | $159 | −40% |
| 30% | $211 | −20% |
| 50% | $264 | 0% |
SayGM runs cascades at the gateway level, so no routing code is needed. On OpenRouter you write the retry yourself.
04
3. Stop paying for the gateway
The same Sonnet 5 workload costs $222.35 on SayGM: $89 for input and $133.35 for output. That is $41.40 a month, or $497 a year, for a base URL change.
- OpenRouter, fee included
- SayGM
Show data as a table
| OpenRouter, fee included | SayGM | |
|---|---|---|
| Side project (5M in · 1M out) | $21 | $18 |
| Production app (50M in · 15M out) | $264 | $222 |
| High volume (500M in · 100M out) | $2,110 | $1,779 |
The gap is wider on other vendors: about 29 percent on GPT-5.5 and 34 percent on Gemini 3.1 Pro. See all prices.
05
4. Trim output before input
Output tokens cost five times more than input on Claude. At an 80/20 mix, the 20 percent of tokens a model writes are 56 percent of the bill.
- Input tokens (80% of volume)
- Output tokens (20% of volume)
Show data as a table
| Model | Share of bill from input | Share of bill from output |
|---|---|---|
| Claude Sonnet 5 | 44% | 56% |
| Claude Haiku 4.5 | 44% | 56% |
| GPT-5.5 | 40% | 60% |
| Gemini 3.1 Pro Preview | 40% | 60% |
Cutting output by 20 percent removes 3M tokens and saves $31.65 a month. Cutting input by 20 percent removes 10M tokens and saves $21.10. Ways to cut output:
- Set
max_tokensper task, not one global ceiling. - Ask for diffs or edits instead of full rewritten files.
- Return structured output with short field names.
- Drop restated context and closing summaries from system prompts that ask for them.
06
5. Keep the cache warm
Prompt caching discounts repeated prefixes such as system prompts, tool definitions and long documents. The saving depends on the provider's cache pricing and your hit rate, so it is not in the table above.
- Put stable content first and variable content last.
- Keep tool definitions in a fixed order.
- Keep a conversation on one provider. Gateways that spread a conversation across providers lose cache hits. SayGM pins a conversation to one provider.
07
All levers together
Apply three levers at once: SayGM as the gateway, output trimmed to 12M tokens, and a Haiku-first cascade with 30 percent escalation. Haiku costs $97.90 for the full volume. Escalations add 30 percent of the trimmed Sonnet cost of $195.68, or $58.70.
Run the numbers for your own volumes in the calculator. Rates are a snapshot; SayGM's pricing page (opens in a new tab) has the current figures.
Prices in this guide are a October 2026 snapshot and move each epoch.
Current SayGM pricing (opens in a new tab)