Skip to content

Gateway

The Gateway is our customer-facing API. Everything a customer can do goes through it: run a job, rent a VM, deploy a model, buy credits. It is also the source of truth other services ask when they need to know what a customer owns.

Planned direction: this is the current state, not the target

/app is being dismantled. Add no new functionality here. Business logic MUST move to the service that owns it. Any surviving thin layer holds auth and quotas only, and MUST NOT be Python. The Supabase store below goes too, with its data moving to Cloud SQL. See /app and /supabase.

Code /app/src/app/ ⧉
Language Python (FastAPI)
Deployment systemd unit gateway on ubuntu_gateway, behind the NGINX proxy
Public URL https://<env>.api.lyceum.technology, plus https://api.lyceum.technology in production
Client lyceum-cli ⧉, the dashboard, and the VS Code extension

For the Gateway's role inside Iris specifically, see Gateway - Iris integration.

What it owns

The Gateway is a coordinator, not a worker. It holds no GPU, runs no model and prices no usage. What it does own:

  • Authentication and identity. API keys, Supabase sessions, org membership.
  • Lifecycle state. Which executions, VMs and deployments exist, and what state they are in.
  • Fan-out to the components that do the work. Streamer for jobs, Hydra for VMs, Kubernetes for Iris replicas, external providers for serverless inference.
  • Reporting billable usage to Croesus. The Gateway emits events. Croesus decides what they cost.

API surface

Three route prefixes, distinguished by who may call them:

Prefix Caller In OpenAPI
/api/v2/external/… Customers, lyceum-cli, dashboard Yes
/api/v2/internal/… Other Lyceum services No
/api/internal/… Callbacks from Execlet, Streamer, Supabase, Stripe No

Internal routers are registered with include_in_schema=False. This is not cosmetic: the OpenAPI schema drives our published SDKs, so anything visible there becomes a customer-facing API we have to keep.

The external surface is grouped into the following areas, which also form the x-tagGroups in the generated docs:

Area What it covers
Serverless GPUs Python, container and compose executions. GPU selection
Dedicated Inference Iris deployment lifecycle
Serverless Inference Chat, embeddings, image and video generation via external providers
Virtual Machines Hydra-provisioned VMs
Storage S3-compatible file and credential management
Account Auth, API keys, environment variables, quotas
Billing Credits, vouchers, Stripe
Organizations Orgs, members, invites
Platform Pricing, machine types, logs, health

Authentication

The Gateway accepts three kinds of bearer token on the same header, and picks the path by inspecting the token:

  • Lyceum API keys. lk_ followed by 64 hex characters. Stored as a SHA-256 hash with an 8-character prefix kept in the clear for identification. The plaintext key is returned exactly once, at creation.
  • Supabase JWTs. Issued by Supabase Auth for dashboard sessions.
  • Internal service tokens. Shared secrets held by InferenceProxy, MetricServer and DeploymentScaler, used for service-to-service calls that act on a customer's behalf.

What it talks to

flowchart TD
    U[Customer / CLI / Dashboard] --> G[Gateway]
    G --> SB[(Supabase Postgres)]
    G --> IR[(CloudSQL — Iris)]
    G --> ST[Streamer]
    G --> HY[Hydra]
    G --> K8[Kubernetes]
    G --> CR[Croesus]
    G --> RD[(Redis)]
    G --> MQ[RabbitMQ]
    G --> S3[(MinIO / S3)]
    G --> SP[Stripe]
Peer Direction Purpose
Streamer out Submit and abort execution jobs, and Iris replicas on the streamer backend
Hydra out Provision and release bare VMs
Kubernetes out Create Iris replicas directly on the k8s backend
Croesus out Record billable events. Read balances and suspension status
Redis out Seed and invalidate the Iris deployment cache read by InferenceProxy
RabbitMQ both Consume Croesus auto_top_up.disabled events
Execlet, Streamer in Job and replica status callbacks
MetricServer, DeploymentScaler, InferenceProxy in Replica health, scaling and resume requests
Supabase Auth in user_created hook
Stripe in Payment webhooks

Data stores

The Gateway reads two databases, which is a wart worth knowing about:

  • Supabase Postgres. Users, orgs, API keys, executions, VMs. Reached both through PostgREST (supabase-py) and directly through SQLAlchemy.
  • CloudSQL. The Iris tables dedicated_deployments and deployment_replicas.

Environments are separated by Postgres schema, not by database: public is production, with integration and development alongside it. Because two access paths each carry their own schema setting, a validate_schema_environment guard runs before anything else at startup and refuses to boot when:

  • a production schema is configured outside the production environment. This holds even if the environment marker is missing entirely.
  • the PostgREST and SQLAlchemy paths disagree on which schema to use.

Warning

This guard is the only thing standing between a misconfigured devel deployment and production customer data. Do not weaken it to make a local setup easier. Set DB_POSTGRES_SCHEMA to a local schema instead.

Background services

Startup launches long-running tasks alongside the request handlers. Each is wrapped so that a failure logs and degrades that feature rather than preventing the Gateway from serving:

Task Purpose
Billing service Meters running models, enforces VM credit suspension, and runs health checks
VM status sync Reconciles VM state against Hydra
Execution queue sync Reconciles execution state against Streamer. Disabled by DISABLE_EXECUTION_QUEUE_SYNC
Hardware notifier Notifies customers waiting on a hardware profile
Auto-top-up consumer Consumes Croesus auto_top_up.disabled events and emails the customer
Serverless inference client Holds the pooled connection to the external provider

Because these run in-process, the Gateway is not horizontally scalable as it stands: a second instance would double-meter models and duplicate the VM suspension checks.

Iris backends

IRIS_BACKEND selects how a dedicated inference replica is materialised:

  • streamer: the Gateway renders a docker-compose job and submits it to Streamer, which schedules it onto an Execlet node.
  • k8s: the Gateway creates a Kubernetes StatefulSet directly. With K8S_AUTOSCALING_ENABLED, a KEDA ScaledObject owns the replica count of a dedicated deployment between its min_replicas and max_replicas.

The k8s path authenticates in one of two ways. A kubeconfig file covers local development and the kind end-to-end harness. GCP OIDC covers the rest: the Gateway mints a Google ID token that the API server maps to a service account. The kubeconfig path takes precedence when set.

Deployment

The Gateway is published as a versioned wheel and installed by the gateway Ansible role. That role templates the environment file and the systemd unit, and restarts the service when either changes.

Warning

The version is recorded in two places and both must be bumped together: app/pyproject.toml and infra/ansible/playbooks/roles/gateway/meta/main.yml. CODEOWNERS carries a dedicated entry for the role file purely because of this. Bumping only the first deploys nothing.

The wheel also vendors shared_models/ ⧉ through a force-include in app/pyproject.toml. A change to a shared model therefore reaches the Gateway only once that wheel is rebuilt and the version bumped.

See Mise tasks for building and testing, and Deploying services for the deployment itself.

Known rough edges

  • Two databases and two access paths per database. The schema guard exists because of this, not in spite of it.
  • Background tasks live in the API process, which blocks running more than one Gateway instance.
  • Startup and shutdown use the deprecated @app.on_event hooks rather than a lifespan context manager.
  • CORS origins are a hard-coded list in main.py, so a new frontend host means a code change and a release.