Summary
- Frontier models cost 16 to 34 percent less than on OpenRouter once its 5.5 percent top-up fee is counted.
- Claude has the smallest gap at about 16 percent. Gemini has the largest at about 34 percent.
- The gateway runs in an Intel TDX enclave with published attestation. Frontier providers still receive prompts.
- About 65 models against several hundred on OpenRouter. Audit your model list before switching.
- Migration is a base URL and key change for the OpenAI, Anthropic and Gemini SDKs.
01
Verdict
SayGM is an LLM gateway built on Bittensor subnet 28. One API key reaches the current Claude, GPT and Gemini lines plus a curated set of open-weight models. There is no subscription and no fee on credit top-ups. Prices are capped at the model maker's list price and usually land below it.
02
Pricing
OpenRouter passes provider list prices through and charges 5.5 percent on card top-ups. SayGM credits top-ups in full and discounts the token price. To compare like with like, every OpenRouter figure below includes the fee.
- SayGM
- OpenRouter, fee included
Show data as a table
| Model | SayGM | OpenRouter, fee included | Difference |
|---|---|---|---|
| Claude Fable 5.1 | $16.00 | $18.99 | -16% |
| GPT-5.5 | $7.45 | $10.55 | -29% |
| Claude Opus 5 | $8.01 | $9.50 | -16% |
| Claude Opus 5.5 | $6.40 | $7.60 | -16% |
| GPT-5.6 Terra | $2.98 | $4.22 | -29% |
| Gemini 3.1 Pro Preview | $2.78 | $4.22 | -34% |
| Claude Sonnet 5.5 | $3.20 | $3.80 | -16% |
| Claude Sonnet 5 | $3.20 | $3.80 | -16% |
| GPT-6.1 Sol | $2.68 | $3.80 | -29% |
| Gemini 3.5 Flash | $2.08 | $3.16 | -34% |
| Kimi K3 | $1.76 | $2.24 | -21% |
| Claude Haiku 4.5 | $1.60 | $1.90 | -16% |
| GPT-5.6 Luna | $0.298 | $0.422 | -29% |
| GPT-6 Luna | $0.13 | $0.194 | -33% |
| DeepSeek V4.1 Flash | $0.288 | $0.13 | 122% |
| Model | SayGM | OpenRouter | vs OpenRouter | vs list |
|---|---|---|---|---|
| Claude Opus 5 | $4.45 / $22.23 | $5.28 / $26.38 | 16% lower | 11% lower |
| Claude Sonnet 5 | $1.78 / $8.89 | $2.11 / $10.55 | 16% lower | 11% lower |
| GPT-5.5 | $3.75 / $22.50 | $5.28 / $31.65 | 29% lower | 25% lower |
| Gemini 3.1 Pro Preview | $1.40 / $8.40 | $2.11 / $12.66 | 34% lower | 30% lower |
| Kimi K3 | $1.02 / $5.08 | $0.42 / $9.50 | 18% lower | n/a |
| DeepSeek V4.1 Flash | $0.09 / $0.36 | $0.03 / $0.53 | 11% higher | n/a |
What a production workload costs
Percentages hide scale. The table below prices a production workload of 50 million input and 15 million output tokens a month, with no caching discount on either side.
| Model | SayGM | OpenRouter | Difference per year |
|---|---|---|---|
| Claude Opus 5 | $556 | $660 | $1,245 |
| Claude Sonnet 5 | $222 | $264 | $497 |
| GPT-5.5 | $525 | $739 | $2,565 |
| Gemini 3.1 Pro Preview | $196 | $295 | $1,193 |
How the price is set
Capacity providers on subnet 28 bid for the right to serve requests. The winning bids set the per-token price, capped at list. SayGM reports an average discount in the low twenties percent across recent epochs. Against list price, our frontier snapshot averages about 21 percent lower, which matches that claim. Against OpenRouter, with its fee, the average is about 25 percent.
03
Catalogue
SayGM lists about 65 models. That covers every current Claude, GPT and Gemini model plus open-weight picks such as Kimi K3, GLM 5.3 and DeepSeek V4.1 Flash. OpenRouter lists several hundred, including legacy and niche models.
- Frontier models are served through each maker's native API.
- Confidential open models, marked as such, run inside the enclave end to end.
- Cascade models fall back to the next model when a call fails.
- Fusion models combine several models in one call. Neither cascade nor fusion has a direct OpenRouter equivalent.
04
Privacy and attestation
The SayGM gateway runs in an Intel TDX trusted execution environment. The host operating system and the machine's operators cannot read memory inside it. Attestation evidence covering TCB status, measurements and gateway identity is published and checked by a third-party verifier.
Who can read a prompt
| Party | SayGM | OpenRouter |
|---|---|---|
| Gateway operator | No. Sealed in the TDX enclave | Processed under a privacy policy |
| Request logging | Not readable by operators | Opt-in |
| Frontier model provider | Yes, with identity stripped | Yes |
| Open-weight model host | Confidential models: enclave only | Depends on the provider |
Optional guardrails swap personal data out before a prompt reaches a frontier provider and back in on the reply. That reduces exposure. It does not remove it. For prompts that must stay sealed, use a model marked Confidential. The private inference explainer covers the limits in more detail.
05
Developer experience
SayGM exposes each model's native API instead of flattening everything into one OpenAI-shaped surface. Anthropic prompt caching, extended thinking and tool use keep working. A conversation stays pinned to one provider, so cached prefixes keep hitting.
| Client | What changes |
|---|---|
| OpenAI SDK | baseURL and apiKey |
| Anthropic SDK | base_url and api_key |
| Vercel AI SDK | baseURL on createOpenAI or createAnthropic |
| Cursor, Cline | Base URL and key fields in settings |
| Claude Code | ANTHROPIC_BASE_URL and ANTHROPIC_API_KEY |
Rate limits are set per key and can be raised on request. A desktop app and a CLI are available for non-API use. Setup notes for each tool are in the integrations index.
06
Where it falls short
- Catalogue size. About 65 models. If one model you depend on is missing, you need a second gateway for it.
- Open-weight input prices. OpenRouter's cheapest providers for DeepSeek V4.1 Flash and Kimi K3 are cheaper on input tokens. Those providers are not confidential.
- No free tier. OpenRouter has free models with tight limits. SayGM is pay as you go only.
- Moving prices. Epoch pricing is usually lower but harder to forecast to the dollar than a fixed rate card.
07
Who should switch
Switch if
- Most of your spend is on Claude, GPT or Gemini.
- You top up often enough that 5.5 percent adds up.
- You need a privacy story that customers or auditors can verify.
Stay on OpenRouter if
- You depend on a long-tail model SayGM does not list.
- Your workload is bulk open-weight inference on public data.
- You rely on free-tier models for prototyping.
Running both is common: SayGM for frontier traffic, OpenRouter for the long tail. The migration checklist covers the code changes and a fallback pattern.
Prices in this guide are a October 2026 snapshot and move each epoch.
Current SayGM pricing (opens in a new tab)