Skip to content

Current state of architecture and repository

What is wrong in each area of the repository, and where it is going. "Legacy" does not mean "removable". It means code that breaks a rule, or is due to be replaced.

What to do about it: engineering rules. This page uses RFC 2119 levels. A MUST here binds as it does there.

Cross-cutting

Build system, container images, versioning and environments are one programme of work, delivered by parallel change. The image pipeline runs on the new tooling. Production still builds and deploys the old way. See environments.

  • Internal interfaces. hydra-autoscaler's externalgrpc.proto is the only committed schema. Every other contract lives in the calling code. gRPC is used by hydra-autoscaler and by stream_execlet, which is being retired. (SERVICE-7)
  • Service identity. Shared *_SERVICE_TOKEN secrets across five Iris services, where Cloud Run ID tokens belong. (SERVICE-8)
  • Build system. Mise ⧉ monorepo tasks are replacing make. The Mise path works today. The Makefiles keep their bodies until their callers move across, then are deleted.
  • Container images. Three shared bases plus one Dockerfile per language family, pushed to the immutable EU registry in platform-staging. The per-service Dockerfiles that predate them go as their services move across. Every Go binary but stream_execlet's can carry a real version. Open: platform-production as the single source of truth, and production cutting over to it.
  • Where services run. Services deployed as binaries onto VMs publish no artefact at all. install stages them into infra/build_artifacts/bin/amd64/ and Ansible copies from there, so a VM runs whatever the pipeline last built. That is the :latest defect. No lint looks there. (SERVICE-9)
  • Environments. staging has its inventory and its digest pins. Its service units are applied by CI on a merge that changes a pin, the rest by hand.

Repository map

Path Purpose Health
/app Python FastAPI gateway To be dismantled
/daedalus_go Shared Go libraries Good, under-adopted
/db Database migrations Needs periodic collapse
/infra Terraform, Ansible, k8s, tooling Mixed
/iris Inference services Boundaries undocumented, partly misplaced
/people Access management Works. SHOULD drive Workspace from IaC
/supabase Supabase data and functions To be removed
/.github CI/CD workflows and bots Duplicated, partly broken

/app

Python FastAPI gateway. Fronts croesus, four Iris services, RabbitMQ, Redis, MinIO and the execlets, while carrying auth, quotas and billing logic of its own.

SHOULD be dismantled. Business logic moves to the service that owns it. Whatever survives handles edge concerns only, and not in Python. No managed API gateway product (SERVICE-11).

/daedalus_go

Shared Go libraries: AMQP, GCP, HTTP, k8s, metrics, messaging, Redis, SQL and more.

  • Mostly good. Needs more documentation.
  • Use SHOULD be enforced over bespoke reinventions. Many services still solve common problems their own way.

/db

Database migrations.

  • Migrations pile up. This is hoarding. They SHOULD be collapsed into a baseline periodically, or the intended schema becomes unreadable.

/infra

/infra/terraform

  • Terragrunt adoption. Terragrunt manages inter-module dependencies. In use for gcp, not for cloudflare, tailscale and the rest. Extending it removes the hacks that feed GCP information into Cloudflare by hand.
  • Legacy module. _modules/gcp/legacy holds very old, uncleaned Terraform. Some is dead. The rest SHOULD move into Terragrunt units.
  • Address plan. The address plan records every allocated range. A new environment takes its ranges from there (INFRA-6).
  • Tailscale and DNS. Workable. The inbound DNS resolver runs. The complexity people meet is the address plan above, not Tailscale.
  • Tailscale state. One tailnet, with one set of ACLs and settings, serves every environment. So its Terraform state is a single shared production state, not one per environment. It stays in the central bucket until production's state moves, then becomes a platform-production unit.
  • Grafana. Production has a stack, with its state in the central bucket. staging has no Grafana yet. Whether each non-production environment gets its own stack is not decided.

/infra/ansible

Playbooks that manage things Ansible SHOULD NOT own:

  • ubuntu-autoscaler.yml SHOULD be a Cloud Run instance.
  • ubuntu-gateway.yml SHOULD be a Cloud Run instance.
  • ubuntu-minio.yml SHOULD proxy GCP (or another S3 provider) instead of running MinIO.
  • ubuntu-proxy.yml SHOULD use GCP load balancers over NGINX.
  • *consul*.yml is in use, not vestigial. The consul role is included by ubuntu-rabbitmq.yml, ubuntu-gateway.yml and ubuntu-execlet.yml, alongside ubuntu-consul-server.yml, consul-data-wipe.yml and consul_dns_client. It goes when the VM estate goes, not before. It is not the answer to service identity. See the cross-cutting entry.

Further debt:

  • Most logic sits in roles. Some playbooks ignore that and do too much, so responsibility is hard to trace. Not worth cleaning up for playbooks due for deletion.
  • Hydra uses some playbooks to set up worker nodes. Ansible is too slow for that. Something faster SHOULD do it. pyinfra is a candidate.
  • Too many inventories: one Neph generates, inventory.gcp.yml, and the dynamic ones under /infra/inventory. We SHOULD converge on one. An improved GCP inventory is the candidate. It also loosens our dependence on Neph for running Ansible.
  • Implicit "A must run before B" ordering between playbooks is neither documented nor enforced.

/infra/k8s

  • The Flux setup is an experiment. It needs extending and hardening before we can rely on it.

/infra/neph

Neph is the infrastructure CLI: terraform, ansible, github, db, k3s, grafana, instances, ssh, tailscale, networking, billing and secrets access. That is what it SHOULD be.

What is left to move:

  • billing: credit-org, debit-org, reset-org-balance and grant-org-signup-credit mutate real customer balances from the infrastructure CLI. The module docstring calls them "for test purposes". Not infrastructure. They belong with croesus, behind real authorisation.
  • people list-emails: exists twice, as a neph command and as //people:list-emails. One goes.
  • Dead code: commands nobody runs. Establish which by usage, then delete.

Folding /infra/ansible and /infra/k8s invocation into neph is contemplated, not decided.

/infra/services

  • ssh-key-manager and test-dummy are unused.
  • The remaining programs SHOULD live at the repository root like everything else, not here.

/infra/build_artifacts

  • Compiler output. The install tasks stage binaries here, in the flat bin/amd64/ layout Ansible and release-infra-deploy.yml resolve against.
  • Container images go to the EU Artifact Registry. Binaries go nowhere. A VM runs an artefact with no version and no digest. Nothing can name what is deployed, or return to it. See "Where services run" above. This directory disappears as those services move to Cloud Run, or shrinks to a staging area for artefacts published first.

/infra/inventory

  • "Inventory" is a poor name, and clashes with Ansible's own "inventory". Prior suggestions: /intent, /manifest.
  • ssh.yml, consul.yml and vpc.yml are poorly maintained and SHOULD be removed. The SOPS files and supported_gpus.yml stay.

/iris

Inference services: deployment_controller, deployment_scaler, inference_proxy, inference_worker, iris_operator, metric_server, dev_tools.

  • Boundaries are undocumented and contested. More services or fewer is an open disagreement. Neither side can argue from anything written down. No service says why its boundary is where it is. None names a point of contact.
  • The split is lopsided. inference_proxy is roughly six times the size of any other service. So "too many" and "too few" are both true, of different parts.
  • DO record the point of contact and the rationale first (SERVICE-10). Then decide granularity with a date on it.
  • inference_proxy invents its own auth layer. Auth SHOULD move elsewhere.
  • Required functionality still lives in /app. It MUST move completely into Iris services.
  • inference_worker is the one Python service in Iris. That is debt, not precedent.
  • iris_operator, the k8s status controller, SHOULD get a full peer review.
  • metric_server serves its own Prometheus registry on /prom instead of the shared /metrics path (SERVICE-2).

/people

Access management, one directory per person: profile.yml for identity and group membership, plus ssh/ and gpg/ public keys.

This directory stays the source of truth. Workspace is the execution layer, and membership SHOULD be applied from these files by Terraform rather than clicked into the admin console.

Workspace cannot replace it. .sops.yaml binds secret access to personal GPG fingerprints, and Workspace has nowhere to hold them. SSH keys could move to OS Login. GPG cannot, until sops key management changes.

Depending on Google for identity is fine while access is expressed in OIDC, SAML and SCIM and provisioning lives in IaC. The console click is the lock-in, not the vendor.

/supabase

Supabase config, functions, migrations and seeds.

  • SHOULD be removed.
  • Move all data to Cloud SQL, or wherever it fits best.

/.github

  • Re-approve bot. Hymenaeus stays until GitHub preserves approvals across a restack. A restack changes the head SHA, which dismisses approvals even when the diff is unchanged. GitHub's stacked pull requests do not change that. As a Cloud Run webhook it would re-approve instantly instead of polling. Unowned, low priority.
  • Duplicated and broken workflows. What each workflow does, which are legacy and which are broken, is in CI workflows.

Enforcement gaps

Rules with no machine check, and what stands between each and a gate. A rule leaves this list by acquiring a check, not by being quietly relaxed.

Rule Gap Blocker
BUILD-25 Four per-service image workflows (build-legacy-*.yml) remain The consolidation of the image workflows
PYTHON-3 The root [tool.ruff] still has no explicit src Harmless while every invocation goes through py-lint.bash, which cds to the root. Fragile if anything ever calls ruff directly
SHELL-2 .shellcheckrc was planned but never landed. //mise:lint covers mise/scripts/ only. Six files in scripts/ still carry no extension The shell-standardisation change. The renames touch the merge-driver config and local aliases
BUILD-6 Nothing checks that config_roots and the per-project mise.toml files agree A small check. Not written
BUILD-7 Nothing rejects a new Makefile Lands when the Makefiles are deleted
BUILD-15 No :latest lint exists production and devel MUST move off floating references first
INFRA-1 Applies are operator-driven, except staging's service units, which deploy-staging.yml applies on a merge CI-driven apply for every other unit and environment. Staging's promotion is the first instance
INFRA-2 cloudflare, tailscale and others are not Terragrunt units Needs a stated deadline and an owner
INFRA-3 No scheduled plan anywhere Not built
INFRA-7 devel and production keep state in the deprecated lyceum-terraform-state, where each keyed environment's account holds roles/storage.admin on the whole bucket. So do the shared Tailscale state and production's Grafana. Their legacy tiers use tier-first prefixes such as cloudflare/production/ production is moving to platform-production, and Tailscale moves with it. devel moves to platform-staging. The legacy tiers move as they become Terragrunt units (INFRA-2)
SERVICE-7 hydra-autoscaler's externalgrpc.proto is the only committed schema. No schema-diff check exists The schema-diff check. Not written
SECURITY-1 Environments other than staging still hold a deploy service-account key in their sops inventory Each environment MUST move to the keyless neph path staging already uses
DOC-6 Nothing tells a reviewer which page a change should have updated The code-path-to-page map. Not written
DB-3 Nothing collapses the migration history Nobody owns the baseline collapse
DOC-2 Nothing measures sentence length or rejects a clause-joining em-dash, semicolon or colon A docs linter. Not written
DOC-4 Nothing checks heading, title, nav label or link-text case The same docs linter
DOC-8 Nothing checks that every page has a nav entry The same docs linter. MkDocs already reports it as INFO, so this is the cheapest of the three