Alerting
Our monitoring setup collects metrics and logs, but collecting them only helps if somebody is looking. Alerting is the part that does the looking for us: Grafana-managed alert rules ⧉ evaluate our metrics inside each Grafana Cloud stack and push a message to Slack when something is wrong.
The concrete rules, routing timers and file locations are listed in the alerting reference. This page explains how the pieces fit together and why.
From metric to Slack message
An alert travels through four stages, and it is worth knowing which one you are debugging when something does not behave:
- Evaluation: Rules are organised into rule groups. Every group has an interval, and all rules in it are evaluated on that interval, sequentially.
- State: A rule whose condition is met does not fire immediately. It first
becomes pending, and only after it has stayed breached for its
forduration does it become firing. This is what keeps a brief spike from paging someone and making noise. - Routing: A firing alert carries its labels to the notification policy, which decides where it goes and how often it is repeated.
- Delivery: The contact point turns it into an actual Slack message (or any other form of contact if set up).
The separation matters because the timers live in different places. How quickly
we notice a problem is the group interval plus for, and it is set per rule.
How often we are reminded about a problem we have already been told about is
repeat_interval, and it is set per route.
Rule is not just single query
It is tempting to write sum(streamer_job_queue_size) > 5 as a single PromQL
query and be done. We deliberately split rules into a chain of stages instead
(query, reduce, and threshold) because a boolean is less useful as an alert. Once
PromQL has collapsed the comparison, all Grafana can tell us is that something
was true. With the threshold as its own stage, the measured value survives into
the alert and we can directly learn how bad the situation actually is.
The chain also gives us somewhere to put the two failure modes that are neither "condition met" nor "condition not met". These are the following: a query returning nothing, and a query that could not run at all.
Routing by label, not by rule
Grafana can attach a contact point directly to a rule. We do not do that. Instead every rule carries labels, and a single notification policy decides what happens to them, because a rule should describe a condition, not a recipient.
The on-call decisions that change far more often than the rules themselves are
who needs to know, how loudly, and how often. Keeping them in one place
means we can change all of them at once. severity is the label that carries
this contract.
The cost is that the root notification policy is a singleton per Grafana stack. Ours replaces whatever default routing Grafana Cloud set up. No second, forgotten path can let a notification escape. Equally, no alert can opt out of it.
Declarative
Alert rules are generated from Python in the repository, exactly like our dashboards. Same reasons: reviewable, diffable and reproducible rather than hand-clicked into a UI.
Production only by default
Unlike dashboards, alerting is not applied to every stack. Only production
gets it. The lower environments spend most of their time in states that look
alarming and are not. Examples: no execlets connected, nothing scheduled, a
service deliberately turned off. Alerts nobody trusts are worse than no alerts.
The tradeoff is that a rule reaches production without ever having fired anywhere. The toggle can be flipped on for a lower environment to try one out.
Provisioned rules are read-only
Because Terraform provisions the rules, Grafana marks them as such and refuses to let anyone edit them in the UI. Tuning an alert is therefore a pull request, not a GUI slider.
What the UI still allows is silencing. Muting a noisy alert during a known maintenance window does not require a deploy.
Synthetic checks
Metric-based alerting can only tell us about things our services report about themselves. It cannot tell us whether a customer can actually run a job, because every component might look healthy while the path through them is broken.
That gap is covered by a scheduled smoke test that submits a real job through the public API every hour and reports to Slack. It predates our Grafana alerting and works completely differently as a GitHub Actions workflow with its own webhook, not a rule in a Grafana stack. This is also why it is the one alert that keeps working when Grafana Cloud is down.