Current state of architecture and repository
What is wrong in each area of the repository, and where it is going. "Legacy" does not mean "removable". It means code that breaks a rule, or is due to be replaced.
What to do about it: engineering rules. This page uses RFC 2119 levels. A MUST here binds as it does there.
Cross-cutting
Build system, container images, versioning and environments are one programme of work, delivered by parallel change. The image pipeline runs on the new tooling. Production still builds and deploys the old way. See environments.
- Internal interfaces.
hydra-autoscaler'sexternalgrpc.protois the only committed schema. Every other contract lives in the calling code. gRPC is used byhydra-autoscalerand bystream_execlet, which is being retired. (SERVICE-7) - Service identity. Shared
*_SERVICE_TOKENsecrets across five Iris services, where Cloud Run ID tokens belong. (SERVICE-8) - Build system. Mise ⧉ monorepo tasks are replacing
make. The Mise path works today. The Makefiles keep their bodies until their callers move across, then are deleted. - Container images. Three shared bases plus one Dockerfile per language family, pushed to the
immutable EU registry in
platform-staging. The per-service Dockerfiles that predate them go as their services move across. Every Go binary butstream_execlet's can carry a real version. Open:platform-productionas the single source of truth, and production cutting over to it. - Where services run. Services deployed as binaries onto VMs publish no artefact at all.
installstages them intoinfra/build_artifacts/bin/amd64/and Ansible copies from there, so a VM runs whatever the pipeline last built. That is the:latestdefect. No lint looks there. (SERVICE-9) - Environments.
staginghas its inventory and its digest pins. Its service units are applied by CI on a merge that changes a pin, the rest by hand.
Repository map
| Path | Purpose | Health |
|---|---|---|
/app |
Python FastAPI gateway | To be dismantled |
/daedalus_go |
Shared Go libraries | Good, under-adopted |
/db |
Database migrations | Needs periodic collapse |
/infra |
Terraform, Ansible, k8s, tooling | Mixed |
/iris |
Inference services | Boundaries undocumented, partly misplaced |
/people |
Access management | Works. SHOULD drive Workspace from IaC |
/supabase |
Supabase data and functions | To be removed |
/.github |
CI/CD workflows and bots | Duplicated, partly broken |
/app
Python FastAPI gateway. Fronts croesus, four Iris services, RabbitMQ, Redis, MinIO and the execlets, while carrying auth, quotas and billing logic of its own.
SHOULD be dismantled. Business logic moves to the service that owns it. Whatever survives handles edge concerns only, and not in Python. No managed API gateway product (SERVICE-11).
/daedalus_go
Shared Go libraries: AMQP, GCP, HTTP, k8s, metrics, messaging, Redis, SQL and more.
- Mostly good. Needs more documentation.
- Use SHOULD be enforced over bespoke reinventions. Many services still solve common problems their own way.
/db
Database migrations.
- Migrations pile up. This is hoarding. They SHOULD be collapsed into a baseline periodically, or the intended schema becomes unreadable.
/infra
/infra/terraform
- Terragrunt adoption. Terragrunt manages inter-module dependencies. In use for
gcp, not forcloudflare,tailscaleand the rest. Extending it removes the hacks that feed GCP information into Cloudflare by hand. - Legacy module.
_modules/gcp/legacyholds very old, uncleaned Terraform. Some is dead. The rest SHOULD move into Terragrunt units. - Address plan. The address plan records every allocated range. A new environment takes its ranges from there (INFRA-6).
- Tailscale and DNS. Workable. The inbound DNS resolver runs. The complexity people meet is the address plan above, not Tailscale.
- Tailscale state. One tailnet, with one set of ACLs and settings, serves every environment. So its
Terraform state is a single shared production state, not one per environment. It stays in the central
bucket until production's state moves, then becomes a
platform-productionunit. - Grafana. Production has a stack, with its state in the central bucket.
staginghas no Grafana yet. Whether each non-production environment gets its own stack is not decided.
/infra/ansible
Playbooks that manage things Ansible SHOULD NOT own:
ubuntu-autoscaler.ymlSHOULD be a Cloud Run instance.ubuntu-gateway.ymlSHOULD be a Cloud Run instance.ubuntu-minio.ymlSHOULD proxy GCP (or another S3 provider) instead of running MinIO.ubuntu-proxy.ymlSHOULD use GCP load balancers over NGINX.*consul*.ymlis in use, not vestigial. Theconsulrole is included byubuntu-rabbitmq.yml,ubuntu-gateway.ymlandubuntu-execlet.yml, alongsideubuntu-consul-server.yml,consul-data-wipe.ymlandconsul_dns_client. It goes when the VM estate goes, not before. It is not the answer to service identity. See the cross-cutting entry.
Further debt:
- Most logic sits in roles. Some playbooks ignore that and do too much, so responsibility is hard to trace. Not worth cleaning up for playbooks due for deletion.
- Hydra uses some playbooks to set up worker nodes. Ansible is too slow for that. Something faster SHOULD do it. pyinfra is a candidate.
- Too many inventories: one Neph generates,
inventory.gcp.yml, and the dynamic ones under/infra/inventory. We SHOULD converge on one. An improved GCP inventory is the candidate. It also loosens our dependence on Neph for running Ansible. - Implicit "A must run before B" ordering between playbooks is neither documented nor enforced.
/infra/k8s
- The Flux setup is an experiment. It needs extending and hardening before we can rely on it.
/infra/neph
Neph is the infrastructure CLI: terraform, ansible, github, db, k3s, grafana,
instances, ssh, tailscale, networking, billing and secrets access. That is what it SHOULD
be.
What is left to move:
billing:credit-org,debit-org,reset-org-balanceandgrant-org-signup-creditmutate real customer balances from the infrastructure CLI. The module docstring calls them "for test purposes". Not infrastructure. They belong with croesus, behind real authorisation.people list-emails: exists twice, as a neph command and as//people:list-emails. One goes.- Dead code: commands nobody runs. Establish which by usage, then delete.
Folding /infra/ansible and /infra/k8s invocation into neph is contemplated, not decided.
/infra/services
ssh-key-managerandtest-dummyare unused.- The remaining programs SHOULD live at the repository root like everything else, not here.
/infra/build_artifacts
- Compiler output. The
installtasks stage binaries here, in the flatbin/amd64/layout Ansible andrelease-infra-deploy.ymlresolve against. - Container images go to the EU Artifact Registry. Binaries go nowhere. A VM runs an artefact with no version and no digest. Nothing can name what is deployed, or return to it. See "Where services run" above. This directory disappears as those services move to Cloud Run, or shrinks to a staging area for artefacts published first.
/infra/inventory
- "Inventory" is a poor name, and clashes with Ansible's own "inventory". Prior suggestions:
/intent,/manifest. ssh.yml,consul.ymlandvpc.ymlare poorly maintained and SHOULD be removed. The SOPS files andsupported_gpus.ymlstay.
/iris
Inference services: deployment_controller, deployment_scaler, inference_proxy,
inference_worker, iris_operator, metric_server, dev_tools.
- Boundaries are undocumented and contested. More services or fewer is an open disagreement. Neither side can argue from anything written down. No service says why its boundary is where it is. None names a point of contact.
- The split is lopsided.
inference_proxyis roughly six times the size of any other service. So "too many" and "too few" are both true, of different parts. - DO record the point of contact and the rationale first (SERVICE-10). Then decide granularity with a date on it.
inference_proxyinvents its own auth layer. Auth SHOULD move elsewhere.- Required functionality still lives in
/app. It MUST move completely into Iris services. inference_workeris the one Python service in Iris. That is debt, not precedent.iris_operator, the k8s status controller, SHOULD get a full peer review.metric_serverserves its own Prometheus registry on/prominstead of the shared/metricspath (SERVICE-2).
/people
Access management, one directory per person: profile.yml for identity and group membership, plus
ssh/ and gpg/ public keys.
This directory stays the source of truth. Workspace is the execution layer, and membership SHOULD be applied from these files by Terraform rather than clicked into the admin console.
Workspace cannot replace it. .sops.yaml binds secret access to personal GPG fingerprints, and
Workspace has nowhere to hold them. SSH keys could move to OS Login. GPG cannot, until sops key
management changes.
Depending on Google for identity is fine while access is expressed in OIDC, SAML and SCIM and provisioning lives in IaC. The console click is the lock-in, not the vendor.
/supabase
Supabase config, functions, migrations and seeds.
- SHOULD be removed.
- Move all data to Cloud SQL, or wherever it fits best.
/.github
- Re-approve bot. Hymenaeus stays until GitHub preserves approvals across a restack. A restack changes the head SHA, which dismisses approvals even when the diff is unchanged. GitHub's stacked pull requests do not change that. As a Cloud Run webhook it would re-approve instantly instead of polling. Unowned, low priority.
- Duplicated and broken workflows. What each workflow does, which are legacy and which are broken, is in CI workflows.
Enforcement gaps
Rules with no machine check, and what stands between each and a gate. A rule leaves this list by acquiring a check, not by being quietly relaxed.
| Rule | Gap | Blocker |
|---|---|---|
| BUILD-25 | Four per-service image workflows (build-legacy-*.yml) remain |
The consolidation of the image workflows |
| PYTHON-3 | The root [tool.ruff] still has no explicit src |
Harmless while every invocation goes through py-lint.bash, which cds to the root. Fragile if anything ever calls ruff directly |
| SHELL-2 | .shellcheckrc was planned but never landed. //mise:lint covers mise/scripts/ only. Six files in scripts/ still carry no extension |
The shell-standardisation change. The renames touch the merge-driver config and local aliases |
| BUILD-6 | Nothing checks that config_roots and the per-project mise.toml files agree |
A small check. Not written |
| BUILD-7 | Nothing rejects a new Makefile |
Lands when the Makefiles are deleted |
| BUILD-15 | No :latest lint exists |
production and devel MUST move off floating references first |
| INFRA-1 | Applies are operator-driven, except staging's service units, which deploy-staging.yml applies on a merge |
CI-driven apply for every other unit and environment. Staging's promotion is the first instance |
| INFRA-2 | cloudflare, tailscale and others are not Terragrunt units |
Needs a stated deadline and an owner |
| INFRA-3 | No scheduled plan anywhere | Not built |
| INFRA-7 | devel and production keep state in the deprecated lyceum-terraform-state, where each keyed environment's account holds roles/storage.admin on the whole bucket. So do the shared Tailscale state and production's Grafana. Their legacy tiers use tier-first prefixes such as cloudflare/production/ |
production is moving to platform-production, and Tailscale moves with it. devel moves to platform-staging. The legacy tiers move as they become Terragrunt units (INFRA-2) |
| SERVICE-7 | hydra-autoscaler's externalgrpc.proto is the only committed schema. No schema-diff check exists |
The schema-diff check. Not written |
| SECURITY-1 | Environments other than staging still hold a deploy service-account key in their sops inventory |
Each environment MUST move to the keyless neph path staging already uses |
| DOC-6 | Nothing tells a reviewer which page a change should have updated | The code-path-to-page map. Not written |
| DB-3 | Nothing collapses the migration history | Nobody owns the baseline collapse |
| DOC-2 | Nothing measures sentence length or rejects a clause-joining em-dash, semicolon or colon | A docs linter. Not written |
| DOC-4 | Nothing checks heading, title, nav label or link-text case |
The same docs linter |
| DOC-8 | Nothing checks that every page has a nav entry |
The same docs linter. MkDocs already reports it as INFO, so this is the cheapest of the three |