Adding alerts
Adding an alert means adding a rule to a
rule group in
infra/terraform/grafana/alerts. Alerts are generated from Python. No .tf
file needs to change, because Terraform discovers rule groups by globbing
that directory.
Add rule
- Decide where the rule belongs. Put it in an existing
*.pyfile if it is about the same service, since rules in one group are evaluated on one shared interval. Otherwise create a new file exposing arule_group()function. - Confirm the metric exists and the query returns what you expect. The fastest
way is Grafana's Explore view against the
grafanacloud-promdatasource in any stack. - Write the rule as a
query chain: a PromQL query, a
reduce, and a threshold.
streamer.pyis an example to copy from. - Choose the labels.
severitydecides the routing, so usecriticalonly for something worth interrupting somebody's evening over. Everything else iswarning. - Write the annotations.
summaryis what lands in Slack, so make it say what broke and how badly.descriptionis where the first debugging step goes. - Decide what "no data" means for this rule. A vanished series is usually a
different incident from the one the rule is about. Then
no_data_state("OK")is right, and the missing series deserves its own rule. - Check the generated JSON:
infra/terraform/grafana/alerts/_generate infra/terraform/grafana/alerts/<file>.py | jq - Apply it:
uv run neph terraform --env production --component grafana - Verify in Grafana under Alerts & IRM → Alerting → Alert rules, in the
Lyceum Alertsfolder. The rule shows its current value, which is the quickest way to see whether the threshold is anywhere near reality.
Test the delivery separately from the rule
Use the Test button on the Slack contact point to confirm notifications
arrive. That takes the routing out of the picture, so a rule that never
notifies is either not firing or not matching a route.
Rules only apply to production
Alerting is
disabled outside production.
Applying the grafana component for devel creates nothing. A new rule
goes live the first time it reaches production.
To try a rule out somewhere safer first, set TF_VAR_grafana_alerting_enabled=true
for that environment. That also needs LYC_GRAFANA_SLACK_WEBHOOK_URL_SECRET
to be readable there, so point it at a scratch channel rather than
#monitoring.
Silencing alerts
To silence an alert without a deploy, add a silence in Grafana under Alerts & IRM → Alerting → Silences. That is the intended escape hatch for maintenance windows.
Changing Slack channel
The channel is fixed by the incoming webhook, not by Terraform, so pointing alerts somewhere else means replacing the secret.
- Create a new Slack incoming webhook for the target channel.
- Write the secret into
infra/inventory/common/config.ymlasLYC_GRAFANA_SLACK_WEBHOOK_URL_SECRETso that the secret is also available for testing in other envs. To change the channel for a single environment, put it ininfra/inventory/<env>/config.ymlinstead, which overrides the common value. - Check that it resolves:
uv run neph config get --env production LYC_GRAFANA_SLACK_WEBHOOK_URL - Apply the
grafanacomponent.