Skip to content

Latest commit

 

History

History
355 lines (305 loc) · 21.9 KB

File metadata and controls

355 lines (305 loc) · 21.9 KB

Project Overview

OpenGradient TEE-gateway is an LLM routing service designed to run within AWS Nitro Enclave TEE (Trusted Execution Environment). It provides a secure, cryptographically verifiable interface to multiple LLM providers (OpenAI, Anthropic, Google Gemini, xAI Grok) with remote attestation, response signing, and x402v2 micropayment access control. The tee-gateway is a part of the decentralized OpenGradient network providing verifiable inference.

The repo must provide a stable AWS Nitro PCR when the code doesn't change in order to allow anyone to reproduce the PCRs locally by building the image as a way to verify what code we are running and also for 3rd party operators to set up their own tee-gateway nodes with the same PCRs in order to participate in the network.

Project Structure highlighting core files

├── tee_gateway/             # Main application package (Flask/connexion)
│   ├── __main__.py          # App factory, x402 middleware setup, key injection; dev server when run directly
│   ├── wsgi.py              # `application` for gunicorn (what the enclave runs)
│   ├── body_preread.py      # Reads/refuses a paid POST body before x402 can answer 402
│   ├── errors.py            # One error payload + outer status for provider vs gateway failures
│   ├── gunicorn_conf.py     # Production server settings: one gthread worker, fatal worker exit
│   ├── llm_backend.py       # LLM provider routing via LangChain, HTTP client management
│   ├── image_generation.py  # Endpoint-based image gen (/images/generations): request shaping, URL→inline-bytes, signed responses
│   ├── tee_manager.py       # TEE key generation, nitriding registration, response signing
│   ├── web_search.py        # In-enclave web search: Exa client, execution, result formatting
│   ├── model_registry.py    # Model config and per-token pricing
│   ├── definitions.py       # On-chain addresses, network IDs, payment amounts
│   ├── facilitator_api.py   # x402 facilitator API client
│   ├── heartbeat/           # Heartbeat/health monitoring
│   ├── controllers/         # Request handlers (chat, completions, security)
│   ├── models/              # OpenAI-compatible Pydantic models
│   ├── openapi/             # openapi.yaml spec
│   └── test/                # Unit tests
├── scripts/
│   ├── start.sh             # Enclave startup script (nitriding + server)
│   ├── run-enclave.sh       # EC2 host launcher (gvproxy, EIF, key injection)
├── pyproject.toml           # Project metadata and dependencies (managed by uv)
├── Dockerfile               # Multi-stage: nitriding builder + python:3.12-slim-bullseye + uv
├── Makefile
└── measurements.txt         # PCR measurements for the deployed enclave image

Common Commands

# Dependency management (uses uv — https://docs.astral.sh/uv/)
uv sync                      # Install/update dependencies from uv.lock
uv add <package>             # Add a new dependency
uv lock                      # Regenerate lockfile after editing pyproject.toml
# IMPORTANT: uv.lock is baked into the Docker image and affects PCR measurements.
# Only regenerate the lockfile when intentionally changing dependencies.

# Run server locally for development (without TEE)
make test-local              # Runs: uv run python -m tee_gateway (Werkzeug dev server)
make serve                   # Runs the production command the enclave uses (gunicorn)

# Linting and type checking
make lint                    # Run ruff format + ruff check + mypy
make mypy                    # Run mypy type checker only

# Build enclave image
make image                   # Build Docker image as TAR using Kaniko

# Build EIF and run in Nitro Enclave
make run                     # or: make all

# Clean build artifacts
make clean

# Show all available targets
make help

Environment Variables

API keys (injected at runtime via POST /v1/keys — do NOT bake into the image):

  • OPENAI_API_KEY
  • ANTHROPIC_API_KEY
  • GOOGLE_API_KEY
  • XAI_API_KEY
  • ARK_API_KEY (BytePlus / ByteDance ModelArk; injected as bytedance_api_key)
  • OPENROUTER_API_KEY (OpenRouter; injected as openrouter_api_key)
  • ZAI_API_KEY (Z.ai Model API; injected as zai_api_key)
  • WAVESPEED_API_KEY (WaveSpeed; injected as wavespeed_api_key)
  • EXA_API_KEY (Exa search; injected as exa_api_key) — backs the in-enclave /v1/web_search endpoint, not an LLM provider. Without it the endpoint returns 503 and /health reports web_search_enabled: false.

Server configuration:

  • API_SERVER_PORT (default: 8000)
  • API_SERVER_HOST (default: 0.0.0.0)
  • EVM_PAYMENT_ADDRESS — wallet address to receive x402 payments
  • FACILITATOR_URL — x402 facilitator endpoint

Architecture

Core Flow

  1. TEEKeyManager (tee_manager.py) generates RSA-2048 key pair on startup and registers the public key hash with the nitriding daemon
  2. Incoming requests pass through x402 payment middleware before reaching handlers
  3. Requests are routed to the appropriate LLM provider via LangChain (llm_backend.py)
  4. All responses are signed with RSA-PSS-SHA256 over keccak256(requestHash || outputHash || timestamp)
  5. Clients verify attestation → get public key → verify signatures

Key Components

  • tee_manager.py: RSA key generation, nitriding registration (/enclave/hash), response signing
  • llm_backend.py: LangChain model instantiation, HTTP client management, provider routing from model name
  • model_registry.py: Maps model names to providers and per-token USD pricing (used by dynamic cost calculator)
  • definitions.py: On-chain constants (addresses, network IDs, payment amounts) — configure here for your deployment
  • web_search.py: Exa HTTP client, search execution, and result formatting/citation extraction (serves /v1/web_search)
  • util.py: dynamic_session_cost_calculator converts actual token usage to x402 payment amounts

API Endpoints

Endpoint Purpose
/health Health check (status, version, tee_enabled, web_search_enabled)
/signing-key TEE public key (PEM) and tee_id
/enclave/attestation Nitro attestation document (served by nitriding)
/v1/keys One-time API key injection (POST, loopback-only)
/v1/completions Text completion (signed)
/v1/chat/completions Chat completion with tool support (signed)
/v1/web_search In-enclave Exa web search (signed, flat per-search price)

TEE Integration

  • Nitriding daemon runs on localhost:8080, provides TLS termination (port 443 externally)
  • Endpoints /enclave/ready and /enclave/hash used for nitriding registration
  • PCR measurements in measurements.txt fingerprint the exact enclave image

Supported Providers

model_registry.py is the source of truth; this list mirrors it. Model name prefixes determine routing:

  • OpenAI: gpt-6-sol, gpt-6-luna, gpt-6-astra, gpt-4.1, gpt-4.1-mini, gpt-5, gpt-5-mini, gpt-5.2, gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, gpt-5.5, gpt-5.6-sol/terra/luna, o3; image generation: gpt-image-2.5-flare, gpt-image-2.5-sunburst, gpt-image-2
  • Anthropic: claude-sonnet-4-5/4-6, claude-sonnet-5, claude-sonnet-5-5, claude-haiku-5-5, claude-haiku-4-5, claude-opus-4-5/4-6/4-7/4-8, claude-opus-5, claude-opus-5-5, claude-fable-5, claude-fable-5-1
  • Google: gemini-3.8-flash, gemini-3.7-flash, gemini-3.6-flash, gemini-3.5-flash, gemini-3.5-flash-lite, gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.5-pro, gemini-3-flash-preview, gemini-3.1-pro-preview; image generation: gemini-2.5-flash-image, gemini-3.1-flash-image, gemini-nano-banana-2.1
  • xAI: grok-4.7, grok-4.6, grok-4.5, grok-4.3, grok-4.20-reasoning, grok-4.20-non-reasoning; image generation: grok-2-image, grok-imagine-image-2.0
  • ByteDance (BytePlus ModelArk, OpenAI-compatible, ap-southeast): seed-1.6, seed-1.8, seed-2.0-lite, deepseek-v4-flash, deepseek-v4-pro, glm-5.2 (Z.ai's model served via a ModelArk deployment endpoint); image generation: seedream-4.0, seedream-5.0-lite, seedance-4.5, seedance-5.0
  • OpenRouter (OpenAI-compatible): hy4-preview, hermes-4-405b, hy3
  • Z.ai (Model API, OpenAI-compatible): image generation: glm-image (glm-5.2 chat is routed through BytePlus ModelArk, see ByteDance above)
  • WaveSpeed (async prediction API): image generation: qwen-image-3.0-pro (Alibaba's model; text-to-image and edit are separate WaveSpeed endpoints)

Models kept registered but no longer offered to new clients — each still resolves so older SDK versions keep working, and each is priced at what the provider actually bills, not at its pre-retirement rate:

  • Retired by xAI on 2026-05-15: grok-4, grok-4-fast, grok-4-1-fast, grok-4-1-fast-non-reasoning, grok-code-fast-1, grok-3, grok-3-mini. xAI still accepts these slugs but silently redirects them (to grok-4.3, or grok-build-0.1 for grok-code-fast-1) and bills every one at grok-4.3's rate.
  • Shutting down at OpenAI on 2026-10-23: gpt-4.1-nano, o4-mini. OpenAI names those exact slugs, so they stop resolving that day; replacements are gpt-5.6-luna and gpt-5.6-terra.
  • Watch 2026-12-11: o3, gpt-5 and gpt-5-mini are undated aliases of snapshots retiring then. OpenAI does not document whether such an alias is repointed or retired with its snapshot; if repointed, they become mispriced the way the xAI slugs were.

A model the provider does not serve at all is removed outright rather than repriced — no request to it can succeed. hermes-4-70b went this way when OpenRouter delisted it. For OpenRouter models, openrouter.ai/api/v1/models is authoritative for availability as well as price; a pricing page keeps quoting a rate after the endpoints are gone.

Image generation via OpenAI (gpt-image-2.5-flare, gpt-image-2.5-sunburst, gpt-image-2), xAI (grok-2-image, grok-imagine-image-2.0), ByteDance (seedream-4.0, seedream-5.0-lite, seedance-4.5, seedance-5.0), and Z.ai (glm-image) is served through a provider /images/generations endpoint rather than the chat path (see image_generation.py), but is surfaced on /v1/chat/completions exactly like Gemini's inline-image models (images returned out-of-band under the message images key). The client always receives inline bytes: providers that hand back a hosted URL (Z.ai, Seedance, Seedream 5.0 Lite) are fetched inside the enclave and inlined as data: URIs (the fetch is guarded: http(s) only, non-public IP hosts rejected, redirects + size capped, and only ever called on provider- response URLs, never client input). Image-to-image editing and multi-image compositing ("add this logo to this photo") send the input images inline (data: URIs / image_url content parts on the latest user turn, up to 10), forwarded to providers that support it. Delivery is one of two per-model paths: ByteDance carries the references inline in the JSON image field of /images/generations; OpenAI gpt-image is routed to its separate /images/edits endpoint, where the references ride as multipart image[] file uploads (only inline data: references are uploaded — a plain-URL reference is skipped rather than dereferenced in the enclave). Per-provider request quirks (response format, n, size/watermark, reference support, edit endpoint) live in model_registry.py. These models are billed a flat per-image price (see per_image_price_usd), not per token. A model whose provider prices by output pixel count can declare resolution tiers (image_resolutions / image_default_resolution): the request's optional resolution field ("1.5K", "2K") picks the tier, which supplies the size keyword, the ratio→pixels table, and the per-image price — omitted means the default tier, and the field is rejected on single-resolution models. Seedream 5.0 (1.5K at $0.045 is the default; 2K at $0.09 is what every request used to be pinned to) and Qwen Image 3.0 Pro (1K $0.04 default, 2K $0.075) are the tiered models today.

WaveSpeed (qwen-image-3.0-pro) is the one image provider that does not answer in a single call. image_generation._generate_wavespeed submits an async prediction to /{model_id} — the …/text-to-image model, or the …/edit model (image_edit_model) when the turn carries references — then polls /predictions/{id}/result (2s, backing off to 5s, 180s deadline → 504, under the relay's 210s read timeout), fetches the output URL like any hosted image, and finally deletes the prediction so WaveSpeed's history no longer holds the prompt or output. WaveSpeed takes references as URLs only: inline data: references are uploaded through its /media/uploads ticket (a keyless PUT to the signed URL, which is a credential and is never logged), plain URLs are passed through, and at most image_max_references (3) are sent. Its input images are billed too — per_reference_image_price_usd ($0.003 each) on top of the per-image price, counted from the references actually forwarded. One image per prediction; n is ignored. Every Qwen request sends enable_prompt_expansion: false (image_extra_params).

Web Search

Web search is a dedicated endpoint — POST /v1/web_search — not a chat feature. It does NOT use any provider's native web search (OpenAI/Anthropic/Google/xAI all have one; those were removed), and the gateway runs no tool loop of its own: the client advertises a web_search function tool to its model, calls this endpoint when the model invokes it, and feeds the returned content back as the tool result. The search runs inside the enclave against Exa (web_search.py), so a query rides the same encrypted OHTTP channel as chat and is never visible to the relay or the gateway operator. Points to keep in mind:

  • The chat/completions web_search request flag is a deprecated no-op. It is still accepted (and still part of the signed request hash when sent) so old clients' requests parse and verify, but it binds nothing and bills nothing.
  • Every text model can search — the tool lives in the client, so this is purely a question of function calling, not of provider search support.
  • Request/response: {"query", "num_results"?, "recency_days"?} in; content (model-ready numbered results), citations (structured sources), and the standard tee_* signing fields out. The request hash covers the canonical (sorted-keys) JSON body; the output hash covers content.
  • Reachable through OHTTP: the inner payload's endpoint field ("web_search") routes the sealed request; absent means chat, so existing OHTTP clients are unaffected. Billing flows through the same outer cost-header / billing-frame channel the relay already consumes.
  • Billing is one flat rate (WEB_SEARCH_PRICE_USD) per search that reached Exa, settled from the response's opengradient block like every paid endpoint. Validation failures (400/503) and Exa failures (502) return no cost block and are never settled. A search that ran but matched nothing IS billed.
  • Exa's self-reported costDollars is diagnostic only; settlement never depends on it.

Content Moderation

Image requests on /v1/chat/completions — image generation, image editing, and inline-image chat models — are scored against OpenAI's free omni-moderation-latest endpoint before any provider is called (moderation.py). The check covers the newest user turn: its prompt text plus any attached images. Plain text chat is not moderated; widening scope there is a deliberate future change (the should_moderate_model predicate is the one gate to widen — everything downstream keys off the flag headers). Points to keep in mind:

  • Fail-open: no OpenAI key or a moderation outage means requests proceed unscored (checked: false); a positive verdict always comes from a real moderation response. /health reports moderation_enabled.
  • Blocking: a request flagged for a category in moderation.BLOCKED_CATEGORIES (default: sexual/minors only) is refused with HTTP 451 + code: "moderation_blocked" and never reaches a provider. All other flagged categories are reported but still served.
  • Response surface: the full verdict (flagged/blocked/categories/scores) rides inside the sealed response body under the moderation key — non-streaming responses and the final SSE frame alike — outside the signed output hash, exactly like images and usage.
  • Relay signal: flagged requests additionally carry content-free X-Moderation-Flagged / X-Moderation-Categories / X-Moderation-Blocked outer headers (forwarded through the OHTTP path) so the relay can run its per-user strike/blacklist policy. Clean traffic carries none of these — it is byte-identical to before.
  • Billing: the moderation call is free and adds no cost block changes; blocked (451) requests produce no opengradient block and are never settled.

Error responses

Every inference endpoint answers failures with one JSON shape, built by tee_gateway/errors.py (non-streaming: the response body; streaming: the terminal in-band SSE data: frame, since the status is already 200):

{"error": "Error code: 529 - {...overloaded_error...}", "exception_type": "APIStatusError",
 "source": "provider", "provider_status": 529, "retryable": true}
  • source is provider when a model provider answered with an error or could not be reached (any exception from a provider SDK or httpx), gateway for a failure inside the enclave. provider_status is the provider's own HTTP status when it returned one; retryable says whether resending the same request has a real chance (provider 5xx/429/overload, resets, timeouts).
  • The outer status follows the same classification: a provider that could not be reached or answered 5xx is 502 (504 for a timeout or 408, 503 for 429), a gateway failure 500. A provider 4xx about the request itself (context too long, bad parameter, unknown model, oversize image) is passed through unchanged: the relay and the browser retry 502/503 once, and resending an invalid request cannot help. The exception is 401/402/403/407 from a provider, which are about the gateway's own key or account and become 502 (a 401 would read as the client's credentials failing; a 402 would be taken by the relay's x402 client for a payment challenge from this gateway). Before this, every failure was a 500, so a browser could not tell an overloaded provider from an enclave bug — and could not distinguish either from the bare 502 nitriding emits when it cannot reach this app at all.
  • Never put prompt or completion text in an error: provider messages describe their refusal, not the user's content, and error bodies are forwarded to the relay as plaintext.

Request bodies are read before any response

tee_gateway/body_preread.py wraps the WSGI stack outside the x402 payment middleware and buffers the body of a POST to a paid route before dispatch. Without it, a 402 challenge on a multi-megabyte body was written while nitriding was still streaming the body in; Go's HTTP server closes a connection with more than 256 KiB of unread body, the reset propagated through gvproxy, and the relay saw either an empty ReadError or a bare 502 instead of the challenge. A body over MAX_PAID_REQUEST_BYTES (20 MiB, the same cap the OHTTP handler enforces) is refused with 413 there, before payment, rather than passed through unread. Keep this the outermost layer: anything that can answer before the body is consumed (payment errors, pricing 503s) must run inside it. The module has no import-time side effects, so it is unit-tested directly (test_body_preread.py).

Production server

The enclave serves the app with gunicorn, one gthread worker (scripts/start.sh → tee_gateway/gunicorn_conf.py, app object in tee_gateway/wsgi.py). python -m tee_gateway still runs Werkzeug's development server for local work only; it used to be what the enclave ran, with one unbounded thread per connection, a 128-entry backlog and Connection: close on every response. The config's constraints are not tunables:

  • workers = 1, no preload. The TEE signing key, the nitriding registration, the injected provider keys and the x402 session store are all state of the one worker process. More workers would mean several signing keys behind one registry entry.
  • An unsolicited worker exit halts gunicorn (child_exit, exit status 1; a normal SIGTERM shutdown still exits 0). A respawned worker would have a new signing key and no provider keys and keep answering requests wrongly. Dying is what the bare process did; keep it that way. scripts/start.sh execs gunicorn so that status is the container's.
  • keepalive = 3600. nitriding's Go proxy pools idle loopback connections with no idle timeout and does not retry a POST it has written; the server must not be the side that closes an idle connection first (that race is a bare 502 to the relay).
  • timeout = 0. The heartbeat watchdog is off because, with worker exit fatal, a false positive under load would take the enclave down.
  • GUNICORN_THREADS caps concurrent requests (each chat stream holds a thread for its duration); API_SERVER_HOST/API_SERVER_PORT bind as before.

Verification Examples

  • examples/verify_attestation.py — Validates AWS Nitro attestation documents against the root CA
  • examples/verify_signature_example.py — Demonstrates request hash and RSA-PSS signature verification

Deployment

Multi-stage Docker build: nitriding compiled from source (brave/nitriding-daemon), then copied into python:3.12.10-slim-bullseye. Dependencies are installed via uv sync --frozen from the lockfile for reproducible builds. scripts/start.sh starts nitriding and then gunicorn (see "Production server"). Enclave launched via scripts/run-enclave.sh with gvproxy as the vsock network bridge, allocating 2 CPUs and 8GB memory.

Port 8000 is forwarded to 127.0.0.1 only on the EC2 host (loopback-only for key injection). Port 443 is forwarded publicly via gvproxy.