Skip to content

Adding alerts

Adding an alert means adding a rule to a rule group in infra/terraform/grafana/alerts. Alerts are generated from Python. No .tf file needs to change, because Terraform discovers rule groups by globbing that directory.

Add rule

  1. Decide where the rule belongs. Put it in an existing *.py file if it is about the same service, since rules in one group are evaluated on one shared interval. Otherwise create a new file exposing a rule_group() function.
  2. Confirm the metric exists and the query returns what you expect. The fastest way is Grafana's Explore view against the grafanacloud-prom datasource in any stack.
  3. Write the rule as a query chain: a PromQL query, a reduce, and a threshold. streamer.py is an example to copy from.
  4. Choose the labels. severity decides the routing, so use critical only for something worth interrupting somebody's evening over. Everything else is warning.
  5. Write the annotations. summary is what lands in Slack, so make it say what broke and how badly. description is where the first debugging step goes.
  6. Decide what "no data" means for this rule. A vanished series is usually a different incident from the one the rule is about. Then no_data_state("OK") is right, and the missing series deserves its own rule.
  7. Check the generated JSON:
    infra/terraform/grafana/alerts/_generate infra/terraform/grafana/alerts/<file>.py | jq
    
  8. Apply it:
    uv run neph terraform --env production --component grafana
    
  9. Verify in Grafana under Alerts & IRM → Alerting → Alert rules, in the Lyceum Alerts folder. The rule shows its current value, which is the quickest way to see whether the threshold is anywhere near reality.

Test the delivery separately from the rule

Use the Test button on the Slack contact point to confirm notifications arrive. That takes the routing out of the picture, so a rule that never notifies is either not firing or not matching a route.

Rules only apply to production

Alerting is disabled outside production. Applying the grafana component for devel creates nothing. A new rule goes live the first time it reaches production.

To try a rule out somewhere safer first, set TF_VAR_grafana_alerting_enabled=true for that environment. That also needs LYC_GRAFANA_SLACK_WEBHOOK_URL_SECRET to be readable there, so point it at a scratch channel rather than #monitoring.

Silencing alerts

To silence an alert without a deploy, add a silence in Grafana under Alerts & IRM → Alerting → Silences. That is the intended escape hatch for maintenance windows.

Changing Slack channel

The channel is fixed by the incoming webhook, not by Terraform, so pointing alerts somewhere else means replacing the secret.

  1. Create a new Slack incoming webhook for the target channel.
  2. Write the secret into infra/inventory/common/config.yml as LYC_GRAFANA_SLACK_WEBHOOK_URL_SECRET so that the secret is also available for testing in other envs. To change the channel for a single environment, put it in infra/inventory/<env>/config.yml instead, which overrides the common value.
  3. Check that it resolves:
    uv run neph config get --env production LYC_GRAFANA_SLACK_WEBHOOK_URL
    
  4. Apply the grafana component.