Skip to content

Environments

What each environment is, where its images come from, and how a version reaches it.

IP ranges: address plan. Rules: delivery and environment rules.

Environments

Where the product runs. Each row is an environment's workload project.

Environment Workload project Purpose Images from
production production-492618 The product US registry, in-project
devel devel-466814 Development US registry, in-project
staging staging-505318 Pre-production, deployed from main platform-staging

staging has its inventory in infra/inventory/staging/ and its pins in infra/terraform/envs/staging.env, every one a digest. CI applies its service units on every push to main. Its other units are applied by hand.

Platform projects

Everything long-lived. Examples: Terraform state, image registries, the CI identity. No workloads, no VPC.

The registries of each platform project serve the environments that use it.

platform-production

Holds the production state of every production system, environment, and service.

Project: platform-production-510011. Serves: production.

A resource every environment uses, but production depends on, keeps its state here as shared production state:

  • GitHub organisation. At platform-production/github/. See GitHub repositories.
  • Image registries. The same europe repositories as platform-staging, from platform/artifact_registries. production's Cloud Run service agent and execlet-worker account read them.
  • CI identity. github-actions, from platform/github, pushes to the registries. Unlike in platform-staging, only workflows on main of lyceum-tech/lyceum may use it, so a branch can never publish a production image.
  • Tailscale. One tailnet, with one set of ACLs and settings, serves every environment, so it has one state, not one per environment. It is still tailscale/terraform.tfstate in the central bucket. It becomes a platform-production unit at platform-production/tailscale/ when production's state moves.

platform-staging

The platform project for all other environments except production. This is where platform changes are proven, before rolling out to platform-production.

Project: platform-staging-506908. Serves: staging, devel.

Why one store rather than one per environment

A per-environment repository conflates two independent things: where the bytes are stored, and which version an environment runs. Immutable versions plus a pin let the second live in each environment's deploy config, which leaves the repository a content-addressable store with no environment dimension. Rebuilding per environment defeats immutability. Copying per environment is machinery that buys nothing.

Consolidating removes a boundary, so the design has to replace it. See BUILD-24 for the per-image repositories, scoped IAM and retention that do. Each store serves only environments at or below its platform project's tier, which is the direction ENV-4 allows.

Registry layout

europe multi-region, host europe-docker.pkg.dev, one repository per image, declared in infra/terraform/_modules/gcp/platform_artifact_registries/variables.tf:

  • Base images: builder, go-runtime, python-runtime
  • Service images: croesus, inference-proxy, inference-worker, deployment-scaler, metric-server, iris-operator, test-dummy, llm-router

Multi-region over single region: the store is production-critical for every environment. Availability outweighs the storage cost, and the local-egress edge europe-west3 would give production.

How a version reaches an environment

An image is built once and promoted by reference. Promotion is a commit changing that environment's pin to a digest already proven in the one before it. No rebuild, no copy (ENV-5).

The pins live in:

  • infra/terraform/envs/<env>.env: image_ref_*, inference_proxy_image, deployment_scaler_image
  • infra/inventory/<env>/config.yml: LYC_GCP_ARTIFACT_REGISTRY_URL and *_IMAGE_REF
  • Kustomize newName fields in infra/k8s/overlays/ and iris/iris_operator/deploy/overlays/
  • Ansible, in infra/ansible/playbooks/k3s-server.yml

Rollback is the same commit backwards.

In staging the merge is the deploy. deploy-staging.yml applies the service units on every push to main. See deploying staging.

Canary and soak

Cloud Run splits traffic by revision tag and weight. A new revision takes a small share, runs long enough to show whether it is healthy, then goes to 100%, or is rolled back. This is how a version is meant to enter an environment.

Not in use yet. Production deploys are manual. staging deploys on merge but sends every revision 100% of the traffic at once. ENV-5 states the model. This section describes the procedure once something runs it.

Sizing a non-production environment

Same services, same topology, same region as production, sized down (ENV-3). Every knob is a per-environment TF_VAR. The expensive production behaviours switch off on their own, because is_production is false. Cloud SQL loses PITR and deletion protection, Memorystore drops to BASIC, and the maintenance track goes to canary.

Cloud SQL availability is not one of them. availability_type is ZONAL in every environment, production included.

That leaves machine types and node counts. Rough europe-west3 figures for where the money is, worth re-checking against the pricing calculator:

TF_VAR_… production non-production ~saving/month
gcp_minio_instance_boot_disk_size 6000 GB 20 GB ~CHF 900
gcp_rabbitmq_num_instances × machine 3 × e2-standard-2 1 × e2-medium ~CHF 130
gcp_default_instance_machine_type n2-standard-4 e2-medium ~CHF 90
gcp_sql_database_main_instance_tier db-custom-2-7680 db-custom-1-3840 ~CHF 60
gcp_k3s_server_num_instances 3 1 ~CHF 50
gcp_k3s_server_instance_machine_type e2-standard-2 e2-medium ~CHF 25
gcp_sql_database_k3s_instance_tier db-custom-2-3840 db-custom-1-3840 ~CHF 30
gcp_tailscale_subnet_router_num_instances 2 1 ~CHF 25

Also set the generic disk and machine knobs small: gcp_default_instance_boot_disk_size=20, lyceum_instance_default_machine_type=e2-medium, gcp_rabbitmq_instance_boot_disk_size=30, gcp_k3s_server_instance_boot_disk_size=50.

GPU workers are the exception. They keep production's autoscaling configuration: an environment that pins them to zero cannot exercise the autoscaler.

What a self-contained environment owns

ENV-1 requires the control plane too, not just the workload services. Two things make that harder than it sounds, and both are why staging is a bigger piece of work than adding an .env file:

  • The API backend. Production serves app from outside this repository's IaC. An environment that does not deploy its own answers <env>.api.lyceum.technology from somewhere else, which is the cross-environment wiring ENV-2 forbids.
  • Supabase. Auth, the data layer, and the schema are Supabase-native: GoTrue, PostgREST, the auth schema, RLS, and the anon/authenticated/service_role roles. LYC_SUPABASE_* lives in inventory/common/config.yml today: one project shared by every environment. So one environment's users, API keys, and billing rows land in the project production uses. Cloud SQL cannot substitute: it is Postgres alone, with no GoTrue and no PostgREST. A dedicated Supabase project gives the isolation at no code cost.

Deploys are keyless. An operator applies with their own ADC, CI with Workload Identity Federation (SECURITY-1). Runtime secrets stay sops-encrypted and are generated fresh per environment (SECURITY-3).

One known exception: staging uses devel's TensorX API key (LYC_TENSORX_API_KEY) until it gets a key of its own. Revoking or rate-limiting that key affects both environments, and their TensorX usage is not told apart. Tracked in LYC-1162.