Monitoring
We observe our services with Grafana Cloud ⧉. We do not run our own Prometheus, Loki and Grafana instances. We push metrics and logs to Grafana Cloud, and explore them in its hosted Grafana. This keeps the storage, retention and availability of our observability data off our plate.
The concrete stacks, endpoints, ports and dashboards are listed in the monitoring reference. This page explains how the pieces fit together and why. What happens once a metric says something is wrong is covered separately in alerting.
One stack per environment
Each deployment environment gets its own Grafana Cloud stack, which gives it an isolated Grafana instance under its own subdomain. Keeping the environments in separate stacks means data from one never mixes with another: each has its own read/write tokens and its own Terraform state.
How data gets there
The agent that ships data to Grafana Cloud is
Grafana Alloy ⧉. It scrapes metrics locally and
forwards them to the stack's hosted Prometheus via remote_write. It also
collects logs and pushes them to the stack's hosted Loki. Both authenticate with
the per-environment tokens.
Alloy runs in two places. On our VMs (execlets, gateways, the autoscaler
host) an Ansible role installs it as a systemd service. On Cloud Run
Terraform deploys it as its own grafana-alloy service, for the services that live there.
What each Alloy instance collects is toggled through the role's variables, so a
plain host and a GPU execlet ship different data. Host metrics come from the
node exporter everywhere. GPU (DCGM) metrics come only from machines with GPUs.
The streamer's and Hydra's service metrics come from wherever those run. For
logs, Alloy tails systemd journal units and Docker Compose containers. A
container opts in and attaches labels through
technology.lyceum.observability.* labels, so its logs arrive in Loki tagged
with its service name.
Declarative dashboards
Our dashboards live in the repository as code, not as hand-edited JSON in the Grafana UI. This keeps them reviewable, diffable and reproducible across every stack.
Each dashboard is a Python script. And because dashboards are regular Python, a change is a normal pull request and gets deployed to every environment the same way.
Being told about problems
Everything above is about making the state of the system visible. Turning that into a message is alerting. That covers alert rules evaluated in the same Grafana stacks, an hourly end-to-end smoke test, and the routing that gets both into Slack.