Skip to content

Alerting

Details for our alerting setup: where rules live, what they check and how they are routed to recipients. See the explanation for how these fit together.

Everything below is provisioned per Grafana Cloud stack by infra/terraform/grafana.

Where things live

Path What
infra/terraform/grafana/alerts/*.py one rule group per file
infra/terraform/grafana/alerts/_generate runner Terraform calls per file
infra/terraform/grafana/alerting.tf contact point, notification policy, translation layer
infra/terraform/grafana/variables.tf grafana_alerting_enabled, grafana_slack_webhook_url

Rules are created in a Lyceum Alerts folder, kept separate from the Lyceum folder that holds dashboards.

Environments

Alerting is only created where it is enabled, which by default is production alone. Dashboards are unaffected and still go to every stack.

grafana_alerting_enabled Effect
unset (default) enabled for production, disabled for devel
true enabled, requires grafana_slack_webhook_url to be set
false disabled, no contact point, policy, folder or rules are created

Rule groups are still generated from Python in every environment even when alerting is disabled, so a broken alert script fails the plan before production.

grafana_slack_webhook_url is only required where alerting is enabled. A variable validation rejects an empty value in that case. Keeping the secret in infra/inventory/production/config.yml rather than the common config therefore works, and is what restricts the webhook to the environment that uses it.

Rule groups

Each *.py file exposes a rule_group() function returning a Grafana Foundation SDK ⧉ RuleGroup builder. All rules in a group share its evaluation interval and are evaluated sequentially.

Query chain

Rules are built from queries referring to each other by ref ID. Expression stages (reduce, threshold, math, resample) are evaluated by Grafana itself and use the __expr__ datasource UID rather than a real datasource.

Ref ID Stage Datasource UID
A PromQL query grafanacloud-prom
B reduce, last over A __expr__
C threshold over B, the rule's condition __expr__

Notification routing via contact point

A single root notification policy routes everything. It is a singleton per Grafana stack and replaces Grafana Cloud's default email routing. It routes alerts to the contact point of the #monitoring ⧉ Slack channel.

Streamer smoke test

An hourly end-to-end check that alerts independently of Grafana, defined in .github/workflows/ops-api-monitoring.yml.

Schedule cron: '16 * * * *' (hourly), plus workflow_dispatch
Concurrency single run at a time (cancel-in-progress: false)
Job script scripts/api_calls_for_monitoring.py --hardware-profile <profile> --wait 3
Pass condition last output line contains completed
Slack script scripts/slack_monitoring_messager.py → SLACK_MONITORING_WEBHOOK_URL
Slack channel #monitoring

A slack-status job runs if: always() after both checks and posts a summary whose wording escalates with the number of failures:

Monitored profiles Message
All pass 🎉 all systems operational
One fails ⚠️ execution issues, @channel alert
Both fail 🚨 both profiles failing, @channel alert

The workflow is additionally marked as failed whenever either profile did not complete, surfacing the failure in GitHub alongside the Slack message.