Skip to content

Monitoring

We observe our services with Grafana Cloud ⧉. We do not run our own Prometheus, Loki and Grafana instances. We push metrics and logs to Grafana Cloud, and explore them in its hosted Grafana. This keeps the storage, retention and availability of our observability data off our plate.

The concrete stacks, endpoints, ports and dashboards are listed in the monitoring reference. This page explains how the pieces fit together and why. What happens once a metric says something is wrong is covered separately in alerting.

One stack per environment

Each deployment environment gets its own Grafana Cloud stack, which gives it an isolated Grafana instance under its own subdomain. Keeping the environments in separate stacks means data from one never mixes with another: each has its own read/write tokens and its own Terraform state.

How data gets there

The agent that ships data to Grafana Cloud is Grafana Alloy ⧉. It scrapes metrics locally and forwards them to the stack's hosted Prometheus via remote_write. It also collects logs and pushes them to the stack's hosted Loki. Both authenticate with the per-environment tokens.

Alloy runs in two places. On our VMs (execlets, gateways, the autoscaler host) an Ansible role installs it as a systemd service. On Cloud Run Terraform deploys it as its own grafana-alloy service, for the services that live there.

What each Alloy instance collects is toggled through the role's variables, so a plain host and a GPU execlet ship different data. Host metrics come from the node exporter everywhere. GPU (DCGM) metrics come only from machines with GPUs. The streamer's and Hydra's service metrics come from wherever those run. For logs, Alloy tails systemd journal units and Docker Compose containers. A container opts in and attaches labels through technology.lyceum.observability.* labels, so its logs arrive in Loki tagged with its service name.

Declarative dashboards

Our dashboards live in the repository as code, not as hand-edited JSON in the Grafana UI. This keeps them reviewable, diffable and reproducible across every stack.

Each dashboard is a Python script. And because dashboards are regular Python, a change is a normal pull request and gets deployed to every environment the same way.

Being told about problems

Everything above is about making the state of the system visible. Turning that into a message is alerting. That covers alert rules evaluated in the same Grafana stacks, an hourly end-to-end smoke test, and the routing that gets both into Slack.