Gateway
The Gateway is our customer-facing API. Everything a customer can do goes through it: run a job, rent a VM, deploy a model, buy credits. It is also the source of truth other services ask when they need to know what a customer owns.
Planned direction: this is the current state, not the target
/app is being dismantled. Add no new functionality here. Business logic MUST move to the service that owns
it. Any surviving thin layer holds auth and quotas only, and MUST NOT be Python. The Supabase store below goes
too, with its data moving to Cloud SQL. See
/app and
/supabase.
| Code | /app/src/app/ ⧉ |
| Language | Python (FastAPI) |
| Deployment | systemd unit gateway on ubuntu_gateway, behind the NGINX proxy |
| Public URL | https://<env>.api.lyceum.technology, plus https://api.lyceum.technology in production |
| Client | lyceum-cli ⧉, the dashboard, and the VS Code extension |
For the Gateway's role inside Iris specifically, see Gateway - Iris integration.
What it owns
The Gateway is a coordinator, not a worker. It holds no GPU, runs no model and prices no usage. What it does own:
- Authentication and identity. API keys, Supabase sessions, org membership.
- Lifecycle state. Which executions, VMs and deployments exist, and what state they are in.
- Fan-out to the components that do the work. Streamer for jobs, Hydra for VMs, Kubernetes for Iris replicas, external providers for serverless inference.
- Reporting billable usage to Croesus. The Gateway emits events. Croesus decides what they cost.
API surface
Three route prefixes, distinguished by who may call them:
| Prefix | Caller | In OpenAPI |
|---|---|---|
/api/v2/external/… |
Customers, lyceum-cli, dashboard |
Yes |
/api/v2/internal/… |
Other Lyceum services | No |
/api/internal/… |
Callbacks from Execlet, Streamer, Supabase, Stripe | No |
Internal routers are registered with include_in_schema=False. This is not cosmetic: the OpenAPI schema drives our
published SDKs, so anything visible there becomes a customer-facing API we have to keep.
The external surface is grouped into the following areas, which also form the
x-tagGroups in the generated docs:
| Area | What it covers |
|---|---|
| Serverless GPUs | Python, container and compose executions. GPU selection |
| Dedicated Inference | Iris deployment lifecycle |
| Serverless Inference | Chat, embeddings, image and video generation via external providers |
| Virtual Machines | Hydra-provisioned VMs |
| Storage | S3-compatible file and credential management |
| Account | Auth, API keys, environment variables, quotas |
| Billing | Credits, vouchers, Stripe |
| Organizations | Orgs, members, invites |
| Platform | Pricing, machine types, logs, health |
Authentication
The Gateway accepts three kinds of bearer token on the same header, and picks the path by inspecting the token:
- Lyceum API keys.
lk_followed by 64 hex characters. Stored as a SHA-256 hash with an 8-character prefix kept in the clear for identification. The plaintext key is returned exactly once, at creation. - Supabase JWTs. Issued by Supabase Auth for dashboard sessions.
- Internal service tokens. Shared secrets held by InferenceProxy, MetricServer and DeploymentScaler, used for service-to-service calls that act on a customer's behalf.
What it talks to
flowchart TD
U[Customer / CLI / Dashboard] --> G[Gateway]
G --> SB[(Supabase Postgres)]
G --> IR[(CloudSQL — Iris)]
G --> ST[Streamer]
G --> HY[Hydra]
G --> K8[Kubernetes]
G --> CR[Croesus]
G --> RD[(Redis)]
G --> MQ[RabbitMQ]
G --> S3[(MinIO / S3)]
G --> SP[Stripe]
| Peer | Direction | Purpose |
|---|---|---|
| Streamer | out | Submit and abort execution jobs, and Iris replicas on the streamer backend |
| Hydra | out | Provision and release bare VMs |
| Kubernetes | out | Create Iris replicas directly on the k8s backend |
| Croesus | out | Record billable events. Read balances and suspension status |
| Redis | out | Seed and invalidate the Iris deployment cache read by InferenceProxy |
| RabbitMQ | both | Consume Croesus auto_top_up.disabled events |
| Execlet, Streamer | in | Job and replica status callbacks |
| MetricServer, DeploymentScaler, InferenceProxy | in | Replica health, scaling and resume requests |
| Supabase Auth | in | user_created hook |
| Stripe | in | Payment webhooks |
Data stores
The Gateway reads two databases, which is a wart worth knowing about:
- Supabase Postgres. Users, orgs, API keys, executions, VMs. Reached both through PostgREST (
supabase-py) and directly through SQLAlchemy. - CloudSQL. The Iris tables
dedicated_deploymentsanddeployment_replicas.
Environments are separated by Postgres schema, not by database: public is production, with integration and
development alongside it. Because two access paths each carry their own schema setting, a
validate_schema_environment
guard runs before anything else at startup and refuses to boot when:
- a production schema is configured outside the production environment. This holds even if the environment marker is missing entirely.
- the PostgREST and SQLAlchemy paths disagree on which schema to use.
Warning
This guard is the only thing standing between a misconfigured devel deployment and production customer data. Do not
weaken it to make a local setup easier. Set DB_POSTGRES_SCHEMA to a local schema instead.
Background services
Startup launches long-running tasks alongside the request handlers. Each is wrapped so that a failure logs and degrades that feature rather than preventing the Gateway from serving:
| Task | Purpose |
|---|---|
| Billing service | Meters running models, enforces VM credit suspension, and runs health checks |
| VM status sync | Reconciles VM state against Hydra |
| Execution queue sync | Reconciles execution state against Streamer. Disabled by DISABLE_EXECUTION_QUEUE_SYNC |
| Hardware notifier | Notifies customers waiting on a hardware profile |
| Auto-top-up consumer | Consumes Croesus auto_top_up.disabled events and emails the customer |
| Serverless inference client | Holds the pooled connection to the external provider |
Because these run in-process, the Gateway is not horizontally scalable as it stands: a second instance would double-meter models and duplicate the VM suspension checks.
Iris backends
IRIS_BACKEND selects how a dedicated inference replica is materialised:
streamer: the Gateway renders a docker-compose job and submits it to Streamer, which schedules it onto an Execlet node.k8s: the Gateway creates a Kubernetes StatefulSet directly. WithK8S_AUTOSCALING_ENABLED, a KEDA ScaledObject owns the replica count of a dedicated deployment between itsmin_replicasandmax_replicas.
The k8s path authenticates in one of two ways. A kubeconfig file covers local development and the kind end-to-end
harness. GCP OIDC covers the rest: the Gateway mints a Google ID token that the API server maps to a service account.
The kubeconfig path takes precedence when set.
Deployment
The Gateway is published as a versioned wheel and installed by the gateway Ansible role. That role templates the
environment file and the systemd unit, and restarts the service when either changes.
Warning
The version is recorded in two places and both must be bumped together: app/pyproject.toml and
infra/ansible/playbooks/roles/gateway/meta/main.yml. CODEOWNERS carries a dedicated entry for the role file purely
because of this. Bumping only the first deploys nothing.
The wheel also vendors shared_models/ ⧉ through a
force-include in app/pyproject.toml. A change to a shared model therefore reaches the Gateway only once that wheel
is rebuilt and the version bumped.
See Mise tasks for building and testing, and Deploying services for the deployment itself.
Known rough edges
- Two databases and two access paths per database. The schema guard exists because of this, not in spite of it.
- Background tasks live in the API process, which blocks running more than one Gateway instance.
- Startup and shutdown use the deprecated
@app.on_eventhooks rather than a lifespan context manager. - CORS origins are a hard-coded list in
main.py, so a new frontend host means a code change and a release.