OpenGradient TEE-gateway is an LLM routing service designed to run within AWS Nitro Enclave TEE (Trusted Execution Environment). It provides a secure, cryptographically verifiable interface to multiple LLM providers (OpenAI, Anthropic, Google Gemini, xAI Grok) with remote attestation, response signing, and x402v2 micropayment access control. The tee-gateway is a part of the decentralized OpenGradient network providing verifiable inference.
The repo must provide a stable AWS Nitro PCR when the code doesn't change in order to allow anyone to reproduce the PCRs locally by building the image as a way to verify what code we are running and also for 3rd party operators to set up their own tee-gateway nodes with the same PCRs in order to participate in the network.
├── tee_gateway/ # Main application package (Flask/connexion)
│ ├── __main__.py # App factory, x402 middleware setup, key injection; dev server when run directly
│ ├── wsgi.py # `application` for gunicorn (what the enclave runs)
│ ├── body_preread.py # Reads/refuses a paid POST body before x402 can answer 402
│ ├── errors.py # One error payload + outer status for provider vs gateway failures
│ ├── gunicorn_conf.py # Production server settings: one gthread worker, fatal worker exit
│ ├── llm_backend.py # LLM provider routing via LangChain, HTTP client management
│ ├── image_generation.py # Endpoint-based image gen (/images/generations): request shaping, URL→inline-bytes, signed responses
│ ├── tee_manager.py # TEE key generation, nitriding registration, response signing
│ ├── web_search.py # In-enclave web search: Exa client, execution, result formatting
│ ├── model_registry.py # Model config and per-token pricing
│ ├── definitions.py # On-chain addresses, network IDs, payment amounts
│ ├── facilitator_api.py # x402 facilitator API client
│ ├── heartbeat/ # Heartbeat/health monitoring
│ ├── controllers/ # Request handlers (chat, completions, security)
│ ├── models/ # OpenAI-compatible Pydantic models
│ ├── openapi/ # openapi.yaml spec
│ └── test/ # Unit tests
├── scripts/
│ ├── start.sh # Enclave startup script (nitriding + server)
│ ├── run-enclave.sh # EC2 host launcher (gvproxy, EIF, key injection)
├── pyproject.toml # Project metadata and dependencies (managed by uv)
├── Dockerfile # Multi-stage: nitriding builder + python:3.12-slim-bullseye + uv
├── Makefile
└── measurements.txt # PCR measurements for the deployed enclave image
# Dependency management (uses uv — https://docs.astral.sh/uv/)
uv sync # Install/update dependencies from uv.lock
uv add <package> # Add a new dependency
uv lock # Regenerate lockfile after editing pyproject.toml
# IMPORTANT: uv.lock is baked into the Docker image and affects PCR measurements.
# Only regenerate the lockfile when intentionally changing dependencies.
# Run server locally for development (without TEE)
make test-local # Runs: uv run python -m tee_gateway (Werkzeug dev server)
make serve # Runs the production command the enclave uses (gunicorn)
# Linting and type checking
make lint # Run ruff format + ruff check + mypy
make mypy # Run mypy type checker only
# Build enclave image
make image # Build Docker image as TAR using Kaniko
# Build EIF and run in Nitro Enclave
make run # or: make all
# Clean build artifacts
make clean
# Show all available targets
make helpAPI keys (injected at runtime via POST /v1/keys — do NOT bake into the image):
OPENAI_API_KEYANTHROPIC_API_KEYGOOGLE_API_KEYXAI_API_KEYARK_API_KEY(BytePlus / ByteDance ModelArk; injected asbytedance_api_key)OPENROUTER_API_KEY(OpenRouter; injected asopenrouter_api_key)ZAI_API_KEY(Z.ai Model API; injected aszai_api_key)WAVESPEED_API_KEY(WaveSpeed; injected aswavespeed_api_key)EXA_API_KEY(Exa search; injected asexa_api_key) — backs the in-enclave/v1/web_searchendpoint, not an LLM provider. Without it the endpoint returns 503 and/healthreportsweb_search_enabled: false.
Server configuration:
API_SERVER_PORT(default: 8000)API_SERVER_HOST(default: 0.0.0.0)EVM_PAYMENT_ADDRESS— wallet address to receive x402 paymentsFACILITATOR_URL— x402 facilitator endpoint
- TEEKeyManager (
tee_manager.py) generates RSA-2048 key pair on startup and registers the public key hash with the nitriding daemon - Incoming requests pass through x402 payment middleware before reaching handlers
- Requests are routed to the appropriate LLM provider via LangChain (
llm_backend.py) - All responses are signed with RSA-PSS-SHA256 over
keccak256(requestHash || outputHash || timestamp) - Clients verify attestation → get public key → verify signatures
tee_manager.py: RSA key generation, nitriding registration (/enclave/hash), response signingllm_backend.py: LangChain model instantiation, HTTP client management, provider routing from model namemodel_registry.py: Maps model names to providers and per-token USD pricing (used by dynamic cost calculator)definitions.py: On-chain constants (addresses, network IDs, payment amounts) — configure here for your deploymentweb_search.py: Exa HTTP client, search execution, and result formatting/citation extraction (serves/v1/web_search)util.py:dynamic_session_cost_calculatorconverts actual token usage to x402 payment amounts
| Endpoint | Purpose |
|---|---|
/health |
Health check (status, version, tee_enabled, web_search_enabled) |
/signing-key |
TEE public key (PEM) and tee_id |
/enclave/attestation |
Nitro attestation document (served by nitriding) |
/v1/keys |
One-time API key injection (POST, loopback-only) |
/v1/completions |
Text completion (signed) |
/v1/chat/completions |
Chat completion with tool support (signed) |
/v1/web_search |
In-enclave Exa web search (signed, flat per-search price) |
- Nitriding daemon runs on localhost:8080, provides TLS termination (port 443 externally)
- Endpoints
/enclave/readyand/enclave/hashused for nitriding registration - PCR measurements in
measurements.txtfingerprint the exact enclave image
model_registry.py is the source of truth; this list mirrors it. Model name
prefixes determine routing:
- OpenAI: gpt-6-sol, gpt-6-luna, gpt-6-astra, gpt-4.1, gpt-4.1-mini, gpt-5, gpt-5-mini, gpt-5.2, gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, gpt-5.5, gpt-5.6-sol/terra/luna, o3; image generation: gpt-image-2.5-flare, gpt-image-2.5-sunburst, gpt-image-2
- Anthropic: claude-sonnet-4-5/4-6, claude-sonnet-5, claude-sonnet-5-5, claude-haiku-5-5, claude-haiku-4-5, claude-opus-4-5/4-6/4-7/4-8, claude-opus-5, claude-opus-5-5, claude-fable-5, claude-fable-5-1
- Google: gemini-3.8-flash, gemini-3.7-flash, gemini-3.6-flash, gemini-3.5-flash, gemini-3.5-flash-lite, gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.5-pro, gemini-3-flash-preview, gemini-3.1-pro-preview; image generation: gemini-2.5-flash-image, gemini-3.1-flash-image, gemini-nano-banana-2.1
- xAI: grok-4.7, grok-4.6, grok-4.5, grok-4.3, grok-4.20-reasoning, grok-4.20-non-reasoning; image generation: grok-2-image, grok-imagine-image-2.0
- ByteDance (BytePlus ModelArk, OpenAI-compatible, ap-southeast): seed-1.6, seed-1.8, seed-2.0-lite, deepseek-v4-flash, deepseek-v4-pro, glm-5.2 (Z.ai's model served via a ModelArk deployment endpoint); image generation: seedream-4.0, seedream-5.0-lite, seedance-4.5, seedance-5.0
- OpenRouter (OpenAI-compatible): hy4-preview, hermes-4-405b, hy3
- Z.ai (Model API, OpenAI-compatible): image generation: glm-image (glm-5.2 chat is routed through BytePlus ModelArk, see ByteDance above)
- WaveSpeed (async prediction API): image generation: qwen-image-3.0-pro (Alibaba's model; text-to-image and edit are separate WaveSpeed endpoints)
Models kept registered but no longer offered to new clients — each still resolves so older SDK versions keep working, and each is priced at what the provider actually bills, not at its pre-retirement rate:
- Retired by xAI on 2026-05-15: grok-4, grok-4-fast, grok-4-1-fast, grok-4-1-fast-non-reasoning, grok-code-fast-1, grok-3, grok-3-mini. xAI still accepts these slugs but silently redirects them (to grok-4.3, or grok-build-0.1 for grok-code-fast-1) and bills every one at grok-4.3's rate.
- Shutting down at OpenAI on 2026-10-23: gpt-4.1-nano, o4-mini. OpenAI names those exact slugs, so they stop resolving that day; replacements are gpt-5.6-luna and gpt-5.6-terra.
- Watch 2026-12-11: o3, gpt-5 and gpt-5-mini are undated aliases of snapshots retiring then. OpenAI does not document whether such an alias is repointed or retired with its snapshot; if repointed, they become mispriced the way the xAI slugs were.
A model the provider does not serve at all is removed outright rather than
repriced — no request to it can succeed. hermes-4-70b went this way when
OpenRouter delisted it. For OpenRouter models, openrouter.ai/api/v1/models
is authoritative for availability as well as price; a pricing page keeps
quoting a rate after the endpoints are gone.
Image generation via OpenAI (gpt-image-2.5-flare, gpt-image-2.5-sunburst, gpt-image-2), xAI (grok-2-image, grok-imagine-image-2.0), ByteDance
(seedream-4.0, seedream-5.0-lite, seedance-4.5, seedance-5.0), and Z.ai (glm-image) is served
through a provider /images/generations endpoint rather than the chat path (see
image_generation.py), but is surfaced on /v1/chat/completions exactly like
Gemini's inline-image models (images returned out-of-band under the message
images key). The client always receives inline bytes: providers that hand back
a hosted URL (Z.ai, Seedance, Seedream 5.0 Lite) are fetched inside the enclave
and inlined as data: URIs (the fetch is guarded: http(s) only, non-public IP
hosts rejected, redirects + size capped, and only ever called on provider-
response URLs, never client input). Image-to-image editing and multi-image
compositing ("add this logo to this photo") send the input images inline
(data: URIs / image_url content parts on the latest user turn, up to 10),
forwarded to providers that support it. Delivery is one of two per-model paths:
ByteDance carries the references inline in the JSON image field of
/images/generations; OpenAI gpt-image is routed to its separate
/images/edits endpoint, where the references ride as multipart image[] file
uploads (only inline data: references are uploaded — a plain-URL reference is
skipped rather than dereferenced in the enclave). Per-provider request quirks
(response format, n, size/watermark, reference support, edit endpoint) live in
model_registry.py. These models are billed a flat per-image
price (see per_image_price_usd), not per token. A model whose provider
prices by output pixel count can declare resolution tiers
(image_resolutions / image_default_resolution): the request's optional
resolution field ("1.5K", "2K") picks the tier, which supplies the
size keyword, the ratio→pixels table, and the per-image price — omitted
means the default tier, and the field is rejected on single-resolution
models. Seedream 5.0 (1.5K at $0.045 is the default; 2K at $0.09 is what
every request used to be pinned to) and Qwen Image 3.0 Pro (1K $0.04 default,
2K $0.075) are the tiered models today.
WaveSpeed (qwen-image-3.0-pro) is the one image provider that does not
answer in a single call. image_generation._generate_wavespeed submits an
async prediction to /{model_id} — the …/text-to-image model, or the
…/edit model (image_edit_model) when the turn carries references — then
polls /predictions/{id}/result (2s, backing off to 5s, 180s deadline → 504,
under the relay's 210s read timeout),
fetches the output URL like any hosted image, and finally deletes the
prediction so WaveSpeed's history no longer holds the prompt or output.
WaveSpeed takes references as URLs only: inline data: references are
uploaded through its /media/uploads ticket (a keyless PUT to the signed
URL, which is a credential and is never logged), plain URLs are passed
through, and at most image_max_references (3) are sent. Its input images are
billed too — per_reference_image_price_usd ($0.003 each) on top of the
per-image price, counted from the references actually forwarded. One image
per prediction; n is ignored. Every Qwen request sends
enable_prompt_expansion: false (image_extra_params).
Web search is a dedicated endpoint — POST /v1/web_search — not a chat feature.
It does NOT use any provider's native web search (OpenAI/Anthropic/Google/xAI
all have one; those were removed), and the gateway runs no tool loop of its own:
the client advertises a web_search function tool to its model, calls this
endpoint when the model invokes it, and feeds the returned content back as the
tool result. The search runs inside the enclave against Exa (web_search.py),
so a query rides the same encrypted OHTTP channel as chat and is never visible
to the relay or the gateway operator. Points to keep in mind:
- The chat/completions
web_searchrequest flag is a deprecated no-op. It is still accepted (and still part of the signed request hash when sent) so old clients' requests parse and verify, but it binds nothing and bills nothing. - Every text model can search — the tool lives in the client, so this is purely a question of function calling, not of provider search support.
- Request/response:
{"query", "num_results"?, "recency_days"?}in;content(model-ready numbered results),citations(structured sources), and the standardtee_*signing fields out. The request hash covers the canonical (sorted-keys) JSON body; the output hash coverscontent. - Reachable through OHTTP: the inner payload's
endpointfield ("web_search") routes the sealed request; absent means chat, so existing OHTTP clients are unaffected. Billing flows through the same outer cost-header / billing-frame channel the relay already consumes. - Billing is one flat rate (
WEB_SEARCH_PRICE_USD) per search that reached Exa, settled from the response'sopengradientblock like every paid endpoint. Validation failures (400/503) and Exa failures (502) return no cost block and are never settled. A search that ran but matched nothing IS billed. - Exa's self-reported
costDollarsis diagnostic only; settlement never depends on it.
Image requests on /v1/chat/completions — image generation, image editing,
and inline-image chat models — are scored against OpenAI's free
omni-moderation-latest endpoint before any provider is called
(moderation.py). The check covers the newest user turn: its prompt text plus
any attached images. Plain text chat is not moderated; widening scope
there is a deliberate future change (the should_moderate_model predicate is
the one gate to widen — everything downstream keys off the flag headers).
Points to keep in mind:
- Fail-open: no OpenAI key or a moderation outage means requests proceed
unscored (
checked: false); a positive verdict always comes from a real moderation response./healthreportsmoderation_enabled. - Blocking: a request flagged for a category in
moderation.BLOCKED_CATEGORIES(default:sexual/minorsonly) is refused with HTTP 451 +code: "moderation_blocked"and never reaches a provider. All other flagged categories are reported but still served. - Response surface: the full verdict (flagged/blocked/categories/scores)
rides inside the sealed response body under the
moderationkey — non-streaming responses and the final SSE frame alike — outside the signed output hash, exactly likeimagesandusage. - Relay signal: flagged requests additionally carry content-free
X-Moderation-Flagged/X-Moderation-Categories/X-Moderation-Blockedouter headers (forwarded through the OHTTP path) so the relay can run its per-user strike/blacklist policy. Clean traffic carries none of these — it is byte-identical to before. - Billing: the moderation call is free and adds no cost block changes;
blocked (451) requests produce no
opengradientblock and are never settled.
Every inference endpoint answers failures with one JSON shape, built by
tee_gateway/errors.py (non-streaming: the response body; streaming: the
terminal in-band SSE data: frame, since the status is already 200):
{"error": "Error code: 529 - {...overloaded_error...}", "exception_type": "APIStatusError",
"source": "provider", "provider_status": 529, "retryable": true}sourceisproviderwhen a model provider answered with an error or could not be reached (any exception from a provider SDK or httpx),gatewayfor a failure inside the enclave.provider_statusis the provider's own HTTP status when it returned one;retryablesays whether resending the same request has a real chance (provider 5xx/429/overload, resets, timeouts).- The outer status follows the same classification: a provider that could not be reached or answered 5xx is 502 (504 for a timeout or 408, 503 for 429), a gateway failure 500. A provider 4xx about the request itself (context too long, bad parameter, unknown model, oversize image) is passed through unchanged: the relay and the browser retry 502/503 once, and resending an invalid request cannot help. The exception is 401/402/403/407 from a provider, which are about the gateway's own key or account and become 502 (a 401 would read as the client's credentials failing; a 402 would be taken by the relay's x402 client for a payment challenge from this gateway). Before this, every failure was a 500, so a browser could not tell an overloaded provider from an enclave bug — and could not distinguish either from the bare 502 nitriding emits when it cannot reach this app at all.
- Never put prompt or completion text in an error: provider messages describe their refusal, not the user's content, and error bodies are forwarded to the relay as plaintext.
tee_gateway/body_preread.py wraps the WSGI stack outside the x402
payment middleware and buffers the body of a POST to a paid route before
dispatch. Without it, a 402 challenge on a multi-megabyte body was written
while nitriding was still streaming the body in; Go's HTTP server closes a
connection with more than 256 KiB of unread body, the reset propagated
through gvproxy, and the relay saw either an empty ReadError or a bare 502
instead of the challenge. A body over MAX_PAID_REQUEST_BYTES (20 MiB, the
same cap the OHTTP handler enforces) is refused with 413 there, before
payment, rather than passed through unread. Keep this the outermost layer:
anything that can answer before the body is consumed (payment errors,
pricing 503s) must run inside it. The module has no import-time side
effects, so it is unit-tested directly (test_body_preread.py).
The enclave serves the app with gunicorn, one gthread worker
(scripts/start.sh → tee_gateway/gunicorn_conf.py, app object in
tee_gateway/wsgi.py). python -m tee_gateway still runs Werkzeug's
development server for local work only; it used to be what the enclave ran,
with one unbounded thread per connection, a 128-entry backlog and
Connection: close on every response. The config's constraints are not
tunables:
workers = 1, no preload. The TEE signing key, the nitriding registration, the injected provider keys and the x402 session store are all state of the one worker process. More workers would mean several signing keys behind one registry entry.- An unsolicited worker exit halts gunicorn (
child_exit, exit status 1; a normal SIGTERM shutdown still exits 0). A respawned worker would have a new signing key and no provider keys and keep answering requests wrongly. Dying is what the bare process did; keep it that way.scripts/start.shexecs gunicorn so that status is the container's. keepalive = 3600. nitriding's Go proxy pools idle loopback connections with no idle timeout and does not retry a POST it has written; the server must not be the side that closes an idle connection first (that race is a bare 502 to the relay).timeout = 0. The heartbeat watchdog is off because, with worker exit fatal, a false positive under load would take the enclave down.GUNICORN_THREADScaps concurrent requests (each chat stream holds a thread for its duration);API_SERVER_HOST/API_SERVER_PORTbind as before.
examples/verify_attestation.py— Validates AWS Nitro attestation documents against the root CAexamples/verify_signature_example.py— Demonstrates request hash and RSA-PSS signature verification
Multi-stage Docker build: nitriding compiled from source (brave/nitriding-daemon), then copied into python:3.12.10-slim-bullseye. Dependencies are installed via uv sync --frozen from the lockfile for reproducible builds. scripts/start.sh starts nitriding and then gunicorn (see "Production server"). Enclave launched via scripts/run-enclave.sh with gvproxy as the vsock network bridge, allocating 2 CPUs and 8GB memory.
Port 8000 is forwarded to 127.0.0.1 only on the EC2 host (loopback-only for key injection). Port 443 is forwarded publicly via gvproxy.