LLM Proxy — Requirements & Design Summary

Server-side proxy that holds the OpenAI API key and serves a Flutter chat app (Android + iOS). Goal: the key never leaves the server, only the real app can use the proxy, and spend is bounded.

Key insight: a client-generated ID is not an identity. Anyone who pulls the proxy URL from the APK/IPA can mint unlimited IDs. All per-client limits are bypassable until the proxy can verify the caller is the genuine app on a genuine device. App attestation is the foundation everything else rests on.

1. System Overview

Flutter App Android / iOS Play Integrity / App Attest token LLM Proxy 1. Attestation verify → issue JWT 2. Auth + rate limit (user / IP / global) 3. Validate + moderate input 4. Build request (server-owned params) 5. Stream response + log usage OpenAI API Chat + Moderation spend cap set TLS + JWT SSE stream API key Secrets manager (key per env)

2. Request Lifecycle

Attestdevice + app JWTshort-lived Rate limittokens, not calls Validatelength + moderation Call OpenAIserver params Stream + logcost per user Any step can reject with 401 / 429 / 400 — reject early, before spending money.

3. Layered Defense Model

Spend ceiling (OpenAI dashboard cap + proxy kill switch) Global + per-IP + per-version rate limits Per-user token budget (JWT-bound) App attestation — only genuine app builds get a JWT

4. Requirements

4.1 Identity & Anti-abuse

IDRequirementPriority
ID-1Verify Play Integrity (Android) and App Attest / DeviceCheck (iOS) server-side before issuing any session credential.P0
ID-2Issue short-lived JWTs (e.g. 15–60 min) bound to the attested app instance; refresh requires re-attestation or a refresh token.P0
ID-3Never trust a client-chosen ID as identity; derive user ID server-side from the attested credential.P0
ID-4Rate limit on multiple axes: per user, per IP, per app version, and a global cap.P0
ID-5Limit tokens (input + output) per window, not just request count.P0
ID-6Hard daily/monthly spend cap in OpenAI dashboard plus a proxy-side kill switch.P0
ID-7Ban list / shadow-ban for abusive users; reject expired or revoked JWTs.P1

4.2 Request Hygiene

IDRequirementPriority
RQ-1Client sends only user messages. Proxy owns model, system prompt, max_tokens, temperature, tools. Raw bodies are never forwarded.P0
RQ-2Cap input message length and total history tokens; truncate or summarize old turns server-side.P0
RQ-3Run OpenAI moderation on user input; reject or flag policy violations.P1
RQ-4Stream via SSE; enforce request timeouts and max concurrent streams per user.P1
RQ-5Strict schema validation on every endpoint; reject unknown fields.P1

4.3 Infrastructure & Operations

IDRequirementPriority
OP-1OpenAI key in a secrets manager; never in config files, logs, or error messages.P0
OP-2TLS everywhere; consider certificate pinning in the Flutter client.P0
OP-3Separate OpenAI keys per environment (dev / staging / prod).P0
OP-4Log tokens, cost, model, latency per user; alert on spend spikes and anomalous patterns.P1
OP-5Provider abstraction layer so OpenAI can be swapped or fallbacks added.P2
OP-6Health checks, graceful degradation, and a maintenance-mode response the app understands.P2

4.4 Privacy & Compliance

IDRequirementPriority
PR-1Decide conversation retention policy; if stored, encrypt at rest and support deletion.P0
PR-2Publish a privacy policy covering AI processing (required by Apple and Google for AI chat apps).P0
PR-3In-app account deletion flow (mandatory on both stores).P0
PR-4Minimize PII in logs; redact message content from operational logs.P1

5. Design Decisions

Auth flow

  • App → attestation token → POST /auth/attest
  • Proxy verifies with Google/Apple → returns JWT + refresh token
  • All chat calls carry Authorization: Bearer <jwt>

Rate limiting

  • Token-bucket per user in Redis (tokens/min + tokens/day)
  • Sliding-window per IP
  • Global semaphore for concurrent upstream streams

Endpoints

  • POST /auth/attest
  • POST /auth/refresh
  • POST /chat (SSE)
  • GET /usage (user's own quota)
  • DELETE /account

Client payload

  • { conversation_id, message } only
  • No model / prompt / params from client
  • History reconstructed server-side (or capped client history, validated)

6. Implementation Order

  1. Attestation + JWT issuance (nothing else matters until this works)
  2. Spend caps + kill switch
  3. Token-based rate limiting (Redis)
  4. Server-owned request builder + input caps
  5. SSE streaming + timeouts
  6. Usage logging + alerting
  7. Moderation, account deletion, privacy policy
  8. Provider abstraction, cert pinning, polish