Deploying staging
staging is deployed by GitHub Actions. Every push to main applies the service units. This page
states the identity, the task and the workflow. Trying a branch on staging
is the developer's path. Bootstrapping platform images
covers building and pushing images by hand.
The deploy identity
GitHub Actions applies Terraform as github-deploy@staging-505318.iam.gserviceaccount.com. The
gcp/service_accounts unit creates it where TF_VAR_github_deploy_ref is set. staging.env
sets it to refs/heads/main and names the units CI deploys in TF_VAR_github_deploy_units.
- It is reachable through Workload Identity Federation from
lyceum-tech/lyceumon that ref only. - It may do exactly what a promotion needs. These are the following: read the project, update the deployed units' Cloud Run resources, act as their runtime accounts, read the secret versions they manage, read this environment's state, write the deployed units' state.
- It may also replace one secret version: croesus's meters config. That secret holds
croesus/config.yaml, which changes with croesus, so a promotion may change it. A new version rolls a new croesus revision. - It may read the Cloud Run images it deploys. Cloud Run checks that whoever updates a service can
read its new image.
platform/artifact_registriesgrants that in platform-staging, to the accountLYC_REGISTRY_CONSUMER_DEPLOY_ACCOUNTnames ininventory/platform-staging/config.yml. A new Cloud Run image MUST be listed there as well. - Every write grant is on the resource itself.
roles/vieweris the only project-wide grant. - Each deployed unit grants its own resources in its
deploy_access.tf. A new Cloud Run resource, runtime account or secret in a unit MUST be added there. - If a unit changes anything else, then the CI apply fails with a 403. Apply that change by hand first.
- It is separate from
github-actions, which pushes images and runs migrations.
The first apply MUST be by hand, as an operator who may set IAM on the platform-staging state bucket, in this order:
uv run neph terraform --env staging --component gcp/service_accounts
uv run neph terraform --env staging --component gcp/github
mise run //infra/neph:deploy-staging
Adopt existing resources
An apply fails with a 409 already exists where the project holds a resource that the unit's
state does not. The staging state moved buckets once, so a unit may still hold such resources.
Import them, then apply again:
uv run neph terragrunt --env staging --component gcp/iris -- plan -out tfplan
uv run neph terragrunt --env staging --component gcp/iris -- show -json tfplan > plan.json
uv run python scripts/gen-imports.py staging gcp/iris plan.json europe-west3 > imports.sh
bash imports.sh
mise run //infra/neph:deploy-staging
gen-imports.py emits one import per resource the plan creates. Review imports.sh first. A
# MANUAL line needs its id filled in by hand. An import of a resource that does not exist fails,
and the script continues with the next one.
Deploy by hand
mise run //infra/neph:deploy-staging
The task runs neph terraform over the units TF_VAR_github_deploy_units in staging.env
names, in that order. Terraform authenticates with your Application Default Credentials. neph
prompts before each apply. Set CI=1 to apply without prompts.
The task applies no other unit. Apply the VM estate, Cloud SQL and networking with neph terraform
yourself.
Deploy from CI
deploy-staging.yml runs the same task. It triggers on these events:
- Every push to
main. neph skips a unit whose plan shows no change. - A
workflow_dispatchfrommain. A dispatch from another ref fails on purpose.
One deploy runs at a time. A deploy is never cancelled. A newer merge replaces a waiting run, so
staging ends on the latest main.
The pins in staging.env decide which images run. A pin MAY name an image built from any branch.
The Terraform always comes from main.
Every push to main builds every service image with build-service-images.yml. Pinning
stays manual. To promote a service:
- For a
maincommit, wait for its build. For another branch, dispatch the build from it. - On a branch off
main, runmise run //infra/docker:pin-staging-images <commit>. - Open a PR with that change. Merge. The workflow applies it.
mise run //infra/docker:push-staging-images <dir> does steps 1 and 2 from your machine.
push-image skips an image whose commit is already pushed, so a re-run rebuilds only what is missing.
The main Terraform MUST configure what the pinned image needs. A branch image can need an
environment variable that main does not pass yet. Its container then fails to start. Merge the
Terraform change first.
The latest merge wins. A pin stays until another merge moves it. To return a service to main, pin
a main commit's image, or revert the pin PR.
Image retention
Each platform-staging repository keeps its most recent versions (keep_most_recent_count, 50 by
default). Every push to main adds one version per service, so an old pin falls out of that window.
- Cleanup MUST stay in dry-run until a policy keeps every pinned image. A validation on
cleanup_dry_runin_modules/gcp/platform_artifact_registriesrejectsfalse. - Before going live, add a policy that keeps every pinned image. Then remove the validation.