Skip to content
TokenGauge
Menu

Explainer · Privacy

Private AI inference, and how to check it

Most gateways describe their privacy in a policy. SayGM runs its gateway inside an Intel TDX trusted execution environment and publishes the attestation. This page traces a request through that setup, and marks where the protection ends.

Updated 9 Oct 2026 · 6 min read
Fig. 1Path of a request through the SayGM gateway
  1. Your application sends the request over an encrypted connection.
  2. The SayGM gateway runs inside an Intel TDX enclave. Inside the enclave it decrypts the request, strips account identity, optionally swaps personal data, and routes it.
  3. For a confidential open-weight model, inference also runs inside the enclave, so the prompt is never readable outside it.
  4. For a frontier model, the prompt leaves the enclave and reaches Anthropic, OpenAI or Google in plaintext, under that provider's API terms.
  5. An independent verifier checks the enclave's hardware-signed attestation report.
Blue marks the sealed path. Frontier requests leave the enclave; confidential open-weight requests do not.

§ 01

Four layers between the prompt and everyone else

  1. 01

    Encrypted to the gateway

    Requests are encrypted in transit and decrypted only inside the SayGM enclave. The host machine and its operators cannot read them.

  2. 02

    Identity stripped

    Frontier providers receive the prompt but not the account that sent it. Optional guardrails swap personal data out before sending and back in on the reply.

  3. 03

    Confidential open models

    Confidential open-weight models, whose IDs end in -tee, run inside the enclave as well. Prompts are decrypted only there and never touch a general-purpose host.

  4. 04

    Public attestation

    Hardware-signed evidence covering TCB status, measurements and gateway identity is published and checked by a third-party verifier.

§ 02

What is protected at each stage

The two model types share the gateway. They differ only at the inference step.

StageFrontier modelConfidential open model
In transit to the gatewayEncryptedEncrypted
Inside the gatewaySealed in enclaveSealed in enclave
Account identityStrippedStripped
During inferencePlaintext at providerSealed in enclave

§ 03

What a TEE does not do

It does not hide prompts from frontier providers

Claude, GPT and Gemini run on Anthropic, OpenAI and Google hardware. The prompt arrives there in plaintext, under that provider's API terms.

It proves code, not intent

Attestation shows which code runs in the enclave. Whether that code is acceptable is a judgement you or an auditor still makes.

It does not cover your side

Logs, caches and analytics in your own application are outside the enclave and keep whatever they record.

§ 04

Models that never leave the enclave

The sealed -tee variants of open-weight models in this index, with snapshot prices per million tokens.

  • Kimi K3

    The largest discount in the SayGM catalogue: 67.5 percent below Moonshot's list price on the open tier. A sealed -tee variant runs inside a TEE for prompts that must stay confidential.

    $2.85 in$14.25 out

  • DeepSeek V4.1 Flash

    40 percent below DeepSeek's list price on SayGM's open tier. OpenRouter's cheapest provider is cheaper still, but not confidential. The sealed -tee variant runs inside an Intel TDX enclave.

    $0.27 in$1.08 out

§ 05

Private inference questions

What is a trusted execution environment?

A trusted execution environment, or TEE, is a hardware-isolated region of a CPU. Code and data inside it are encrypted in memory and unreadable by the operating system, the hypervisor and the machine's operator. SayGM uses Intel TDX.

What is remote attestation?

Remote attestation is a hardware-signed report of which code is running inside the enclave. SayGM publishes the evidence and a third-party verifier checks it, so the running gateway code can be compared to what is claimed.

Which models are fully confidential?

Confidential open-weight models on SayGM, whose IDs end in -tee (for example the TEE variants of Kimi K3 and DeepSeek V4.1 Flash), run inside the enclave end to end. The same models on the open tier do not. Frontier models from Anthropic, OpenAI and Google run on those providers' hardware under their API terms.

How does this compare to OpenRouter's privacy?

OpenRouter's privacy rests on a published policy with opt-in logging. SayGM's gateway privacy rests on hardware isolation that can be checked through attestation. In both cases frontier prompts reach the model provider.