Environments
What each environment is, where its images come from, and how a version reaches it.
IP ranges: address plan. Rules: delivery and environment rules.
Environments
Where the product runs. Each row is an environment's workload project.
| Environment | Workload project | Purpose | Images from |
|---|---|---|---|
production |
production-492618 |
The product | US registry, in-project |
devel |
devel-466814 |
Development | US registry, in-project |
staging |
staging-505318 |
Pre-production, deployed from main |
platform-staging |
staging has its inventory in infra/inventory/staging/ and its pins in
infra/terraform/envs/staging.env, every one a digest. CI applies its service units on every
push to main. Its other units are applied by hand.
Platform projects
Everything long-lived. Examples: Terraform state, image registries, the CI identity. No workloads, no VPC.
The registries of each platform project serve the environments that use it.
platform-production
Holds the production state of every production system, environment, and service.
Project: platform-production-510011. Serves: production.
A resource every environment uses, but production depends on, keeps its state here as shared production state:
- GitHub organisation. At
platform-production/github/. See GitHub repositories. - Image registries. The same
europerepositories asplatform-staging, fromplatform/artifact_registries.production's Cloud Run service agent andexeclet-workeraccount read them. - CI identity.
github-actions, fromplatform/github, pushes to the registries. Unlike inplatform-staging, only workflows onmainoflyceum-tech/lyceummay use it, so a branch can never publish a production image. - Tailscale. One tailnet, with one set of ACLs and settings, serves every environment, so it has one
state, not one per environment. It is still
tailscale/terraform.tfstatein the central bucket. It becomes aplatform-productionunit atplatform-production/tailscale/when production's state moves.
platform-staging
The platform project for all other environments except production.
This is where platform changes are proven, before rolling out to platform-production.
Project: platform-staging-506908. Serves: staging, devel.
Why one store rather than one per environment
A per-environment repository conflates two independent things: where the bytes are stored, and which version an environment runs. Immutable versions plus a pin let the second live in each environment's deploy config, which leaves the repository a content-addressable store with no environment dimension. Rebuilding per environment defeats immutability. Copying per environment is machinery that buys nothing.
Consolidating removes a boundary, so the design has to replace it. See BUILD-24 for the per-image repositories, scoped IAM and retention that do. Each store serves only environments at or below its platform project's tier, which is the direction ENV-4 allows.
Registry layout
europe multi-region, host europe-docker.pkg.dev, one repository per image, declared in
infra/terraform/_modules/gcp/platform_artifact_registries/variables.tf:
- Base images:
builder,go-runtime,python-runtime - Service images:
croesus,inference-proxy,inference-worker,deployment-scaler,metric-server,iris-operator,test-dummy,llm-router
Multi-region over single region: the store is production-critical for every environment.
Availability outweighs the storage cost, and the local-egress edge europe-west3 would give
production.
How a version reaches an environment
An image is built once and promoted by reference. Promotion is a commit changing that environment's pin to a digest already proven in the one before it. No rebuild, no copy (ENV-5).
The pins live in:
infra/terraform/envs/<env>.env:image_ref_*,inference_proxy_image,deployment_scaler_imageinfra/inventory/<env>/config.yml:LYC_GCP_ARTIFACT_REGISTRY_URLand*_IMAGE_REF- Kustomize
newNamefields ininfra/k8s/overlays/andiris/iris_operator/deploy/overlays/ - Ansible, in
infra/ansible/playbooks/k3s-server.yml
Rollback is the same commit backwards.
In staging the merge is the deploy. deploy-staging.yml applies the service units on
every push to main. See deploying staging.
Canary and soak
Cloud Run splits traffic by revision tag and weight. A new revision takes a small share, runs long enough to show whether it is healthy, then goes to 100%, or is rolled back. This is how a version is meant to enter an environment.
Not in use yet. Production deploys are manual. staging deploys on merge but sends every
revision 100% of the traffic at once. ENV-5 states the model. This
section describes the procedure once something runs it.
Sizing a non-production environment
Same services, same topology, same region as production, sized down
(ENV-3). Every knob is a per-environment TF_VAR. The expensive
production behaviours switch off on their own, because is_production is false. Cloud SQL loses
PITR and deletion protection, Memorystore drops to BASIC, and the maintenance track goes to
canary.
Cloud SQL availability is not one of them. availability_type is ZONAL in every environment,
production included.
That leaves machine types and node counts. Rough europe-west3 figures for where the money is,
worth re-checking against the pricing calculator:
TF_VAR_… |
production | non-production | ~saving/month |
|---|---|---|---|
gcp_minio_instance_boot_disk_size |
6000 GB | 20 GB | ~CHF 900 |
gcp_rabbitmq_num_instances × machine |
3 × e2-standard-2 | 1 × e2-medium | ~CHF 130 |
gcp_default_instance_machine_type |
n2-standard-4 | e2-medium | ~CHF 90 |
gcp_sql_database_main_instance_tier |
db-custom-2-7680 | db-custom-1-3840 | ~CHF 60 |
gcp_k3s_server_num_instances |
3 | 1 | ~CHF 50 |
gcp_k3s_server_instance_machine_type |
e2-standard-2 | e2-medium | ~CHF 25 |
gcp_sql_database_k3s_instance_tier |
db-custom-2-3840 | db-custom-1-3840 | ~CHF 30 |
gcp_tailscale_subnet_router_num_instances |
2 | 1 | ~CHF 25 |
Also set the generic disk and machine knobs small: gcp_default_instance_boot_disk_size=20,
lyceum_instance_default_machine_type=e2-medium, gcp_rabbitmq_instance_boot_disk_size=30,
gcp_k3s_server_instance_boot_disk_size=50.
GPU workers are the exception. They keep production's autoscaling configuration: an environment that pins them to zero cannot exercise the autoscaler.
What a self-contained environment owns
ENV-1 requires the control plane too, not just the workload services.
Two things make that harder than it sounds, and both are why staging is a bigger piece of work
than adding an .env file:
- The API backend. Production serves
appfrom outside this repository's IaC. An environment that does not deploy its own answers<env>.api.lyceum.technologyfrom somewhere else, which is the cross-environment wiring ENV-2 forbids. - Supabase. Auth, the data layer, and the schema are Supabase-native: GoTrue, PostgREST, the
authschema, RLS, and theanon/authenticated/service_roleroles.LYC_SUPABASE_*lives ininventory/common/config.ymltoday: one project shared by every environment. So one environment's users, API keys, and billing rows land in the project production uses. Cloud SQL cannot substitute: it is Postgres alone, with no GoTrue and no PostgREST. A dedicated Supabase project gives the isolation at no code cost.
Deploys are keyless. An operator applies with their own ADC, CI with Workload Identity Federation (SECURITY-1). Runtime secrets stay sops-encrypted and are generated fresh per environment (SECURITY-3).
One known exception: staging uses devel's TensorX API key (LYC_TENSORX_API_KEY) until it
gets a key of its own. Revoking or rate-limiting that key affects both environments, and their
TensorX usage is not told apart. Tracked in LYC-1162.
Related
- Address plan: the IP ranges each environment owns
- Delivery and environment rules
- CI workflows: what builds and pushes the images
- Current state of the architecture