Skip to content

Deploying staging

staging is deployed by GitHub Actions. Every push to main applies the service units. This page states the identity, the task and the workflow. Trying a branch on staging is the developer's path. Bootstrapping platform images covers building and pushing images by hand.

The deploy identity

GitHub Actions applies Terraform as github-deploy@staging-505318.iam.gserviceaccount.com. The gcp/service_accounts unit creates it where TF_VAR_github_deploy_ref is set. staging.env sets it to refs/heads/main and names the units CI deploys in TF_VAR_github_deploy_units.

  • It is reachable through Workload Identity Federation from lyceum-tech/lyceum on that ref only.
  • It may do exactly what a promotion needs. These are the following: read the project, update the deployed units' Cloud Run resources, act as their runtime accounts, read the secret versions they manage, read this environment's state, write the deployed units' state.
  • It may also replace one secret version: croesus's meters config. That secret holds croesus/config.yaml, which changes with croesus, so a promotion may change it. A new version rolls a new croesus revision.
  • It may read the Cloud Run images it deploys. Cloud Run checks that whoever updates a service can read its new image. platform/artifact_registries grants that in platform-staging, to the account LYC_REGISTRY_CONSUMER_DEPLOY_ACCOUNT names in inventory/platform-staging/config.yml. A new Cloud Run image MUST be listed there as well.
  • Every write grant is on the resource itself. roles/viewer is the only project-wide grant.
  • Each deployed unit grants its own resources in its deploy_access.tf. A new Cloud Run resource, runtime account or secret in a unit MUST be added there.
  • If a unit changes anything else, then the CI apply fails with a 403. Apply that change by hand first.
  • It is separate from github-actions, which pushes images and runs migrations.

The first apply MUST be by hand, as an operator who may set IAM on the platform-staging state bucket, in this order:

uv run neph terraform --env staging --component gcp/service_accounts
uv run neph terraform --env staging --component gcp/github
mise run //infra/neph:deploy-staging

Adopt existing resources

An apply fails with a 409 already exists where the project holds a resource that the unit's state does not. The staging state moved buckets once, so a unit may still hold such resources. Import them, then apply again:

uv run neph terragrunt --env staging --component gcp/iris -- plan -out tfplan
uv run neph terragrunt --env staging --component gcp/iris -- show -json tfplan > plan.json
uv run python scripts/gen-imports.py staging gcp/iris plan.json europe-west3 > imports.sh
bash imports.sh
mise run //infra/neph:deploy-staging

gen-imports.py emits one import per resource the plan creates. Review imports.sh first. A # MANUAL line needs its id filled in by hand. An import of a resource that does not exist fails, and the script continues with the next one.

Deploy by hand

mise run //infra/neph:deploy-staging

The task runs neph terraform over the units TF_VAR_github_deploy_units in staging.env names, in that order. Terraform authenticates with your Application Default Credentials. neph prompts before each apply. Set CI=1 to apply without prompts.

The task applies no other unit. Apply the VM estate, Cloud SQL and networking with neph terraform yourself.

Deploy from CI

deploy-staging.yml runs the same task. It triggers on these events:

  • Every push to main. neph skips a unit whose plan shows no change.
  • A workflow_dispatch from main. A dispatch from another ref fails on purpose.

One deploy runs at a time. A deploy is never cancelled. A newer merge replaces a waiting run, so staging ends on the latest main.

The pins in staging.env decide which images run. A pin MAY name an image built from any branch. The Terraform always comes from main.

Every push to main builds every service image with build-service-images.yml. Pinning stays manual. To promote a service:

  1. For a main commit, wait for its build. For another branch, dispatch the build from it.
  2. On a branch off main, run mise run //infra/docker:pin-staging-images <commit>.
  3. Open a PR with that change. Merge. The workflow applies it.

mise run //infra/docker:push-staging-images <dir> does steps 1 and 2 from your machine. push-image skips an image whose commit is already pushed, so a re-run rebuilds only what is missing.

The main Terraform MUST configure what the pinned image needs. A branch image can need an environment variable that main does not pass yet. Its container then fails to start. Merge the Terraform change first.

The latest merge wins. A pin stays until another merge moves it. To return a service to main, pin a main commit's image, or revert the pin PR.

Image retention

Each platform-staging repository keeps its most recent versions (keep_most_recent_count, 50 by default). Every push to main adds one version per service, so an old pin falls out of that window.

  • Cleanup MUST stay in dry-run until a policy keeps every pinned image. A validation on cleanup_dry_run in _modules/gcp/platform_artifact_registries rejects false.
  • Before going live, add a policy that keeps every pinned image. Then remove the validation.