Alerting
Details for our alerting setup: where rules live, what they check and how they are routed to recipients. See the explanation for how these fit together.
Everything below is provisioned per Grafana Cloud
stack by infra/terraform/grafana.
Where things live
| Path | What |
|---|---|
infra/terraform/grafana/alerts/*.py |
one rule group per file |
infra/terraform/grafana/alerts/_generate |
runner Terraform calls per file |
infra/terraform/grafana/alerting.tf |
contact point, notification policy, translation layer |
infra/terraform/grafana/variables.tf |
grafana_alerting_enabled, grafana_slack_webhook_url |
Rules are created in a Lyceum Alerts folder, kept separate from the Lyceum
folder that holds dashboards.
Environments
Alerting is only created where it is enabled, which by default is production
alone. Dashboards are unaffected and still go to every stack.
grafana_alerting_enabled |
Effect |
|---|---|
| unset (default) | enabled for production, disabled for devel |
true |
enabled, requires grafana_slack_webhook_url to be set |
false |
disabled, no contact point, policy, folder or rules are created |
Rule groups are still generated from Python in every environment even when alerting is disabled, so a broken alert script fails the plan before production.
grafana_slack_webhook_url is only required where alerting is enabled. A
variable validation rejects an empty value in that case. Keeping the secret in
infra/inventory/production/config.yml rather than the common config therefore
works, and is what restricts the webhook to the environment that uses it.
Rule groups
Each *.py file exposes a rule_group() function returning a
Grafana Foundation SDK ⧉
RuleGroup builder. All rules in a group share its evaluation interval and are
evaluated sequentially.
Query chain
Rules are built from queries referring to each other by ref ID. Expression
stages (reduce, threshold, math, resample) are evaluated by Grafana
itself and use the __expr__ datasource UID rather than a real datasource.
| Ref ID | Stage | Datasource UID |
|---|---|---|
A |
PromQL query | grafanacloud-prom |
B |
reduce, last over A |
__expr__ |
C |
threshold over B, the rule's condition |
__expr__ |
Notification routing via contact point
A single root notification policy routes everything. It is a singleton per Grafana stack and replaces Grafana Cloud's default email routing. It routes alerts to the contact point of the #monitoring ⧉ Slack channel.
Streamer smoke test
An hourly end-to-end check that alerts independently of Grafana, defined in
.github/workflows/ops-api-monitoring.yml.
| Schedule | cron: '16 * * * *' (hourly), plus workflow_dispatch |
| Concurrency | single run at a time (cancel-in-progress: false) |
| Job script | scripts/api_calls_for_monitoring.py --hardware-profile <profile> --wait 3 |
| Pass condition | last output line contains completed |
| Slack script | scripts/slack_monitoring_messager.py → SLACK_MONITORING_WEBHOOK_URL |
| Slack channel | #monitoring |
A slack-status job runs if: always() after both checks and posts a summary
whose wording escalates with the number of failures:
| Monitored profiles | Message |
|---|---|
| All pass | 🎉 all systems operational |
| One fails | ⚠️ execution issues, @channel alert |
| Both fail | 🚨 both profiles failing, @channel alert |
The workflow is additionally marked as failed whenever either profile did not complete, surfacing the failure in GitHub alongside the Slack message.