Skip to content

fix: alert on every monitor outage, and resolve the page when a snoozed rule's monitor recovers - #394

Open
FrameAutomata wants to merge 1 commit into
mainfrom
fix/387-391-monitor-alert-cooldown-and-snooze
Open

FrameAutomata wants to merge 1 commit into
mainfrom
fix/387-391-monitor-alert-cooldown-and-snooze

Conversation

@FrameAutomata

@FrameAutomata FrameAutomata commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #387
Closes #391

Two defects in OnCheckStateChange (backend/app/notifications/check_state.go), the notifier behind Monitor Down rules. They sit a few lines apart in one loop, so they are fixed together.

Problem

#387: a second outage inside the rule's cooldown was never alerted. ProcessOutcome calls the notifier only on a state change, so an outage produces exactly one "went down" transition. The notifier dropped that transition when the rule's dedup key for the monitor had been recorded within CooldownMinutes (default 15), and nothing fired again while the monitor stayed down. The alert was not delayed, it was lost; on an escalation channel no page opened. The recovery was not deduplicated, so the channel still got a "recovered" notice for an outage nobody was told about.

#391: a page stayed open when its monitor recovered while the rule was snoozed. The snooze check ran before both the "went down" and the recovery handling, so it skipped the page auto-resolve along with the notifications. A recovery produces a single transition, so nothing resolved the page later.

Fix

  • On recovery, forget the rule's dedup key for that monitor (dedupTracker.forget). The cooldown then covers one outage, which is how the docs already describe it ("prevents a rule from firing repeatedly for the same ongoing condition"): a recovery ends the condition, and the next outage is a new one.
  • Run the two things that end an outage, the dedup reset and the escalation auto-resolve, before the snooze check. Ending an outage is not a notification. A snoozed rule still sends nothing: no alert, no page, no recovery notice.

Behaviour change to be aware of

A flapping monitor now alerts for every outage. Before, it sent the first alert and then only recoveries until the cooldown ran out, which is neither quiet nor informative. The damping for a flapping check is the monitor's failure threshold (consecutive failures before it counts as down), and the docs now say so.

On an escalation channel this means each outage opens its own page, after the previous one auto-resolved.

Verification

Tests, written first and run against unmodified production code, where the three regression tests failed on their assertions:

Test On main
TestCheckDownAlertsForEveryOutage (Slack channel: down, up, down, up) queued down, recovered, recovered; wants down, recovered, down, recovered
TestSecondOutageInsideCooldownOpensNewPage (escalation channel) expected the second outage to open a page
TestRecoveryResolvesPageWhileRuleSnoozed (escalation channel) page status = open, want resolved

Two guards pass on main and on this branch: TestCheckDownCooldownHoldsWithinOneOutage (a repeated down transition with no recovery between still alerts once) and TestSnoozedCheckDownRuleSendsNothing. The existing TestCheckDownOpensPageAndRecoveryAutoResolves now shares a seeding helper with the two new on-call tests.

go test for notifications, oncall, outbox, synthetics and controllers, go vet ./app/... and gofmt -l are clean on the default build; go vet on the two touched packages is clean under telemetry_duckdb and transactional_pg telemetry_ch.

Live, the same script against a build of main (895b03c1) and of this branch, default SQLite build: one TCP monitor (failureThreshold: 1) feeding a Slack rule and an on-call rule, both on the default 15 minute cooldown. The on-call rule is snoozed during the first outage and un-snoozed after it.

Step main This branch
Outage 1 begins Slack alert, page #1 open Slack alert, page #1 open
On-call rule snoozed, outage 1 ends Slack recovery, page #1 still open Slack recovery, page #1 resolved
Rule un-snoozed, outage 2 begins (under a minute after the first) no Slack alert, no new page Slack alert, page #2 open
Outage 2 ends Slack recovery Slack recovery, page #2 resolved
Slack messages in total 3: down, recovered, recovered 4: down, recovered, down, recovered

The monitor recorded 2 incidents in both runs.

CI note

Backend Vulncheck will fail on this PR: main fails that gate on GO-2026-6505 in a dependency, which #386 fixes. This branch does not carry the bump, so the gate clears once #386 has merged and the ci label is re-applied.

🤖 Generated with Claude Code

…ed rule's monitor recovers

Two defects in OnCheckStateChange, the notifier behind check_down rules.

A second outage inside the rule's cooldown was never alerted. The notifier
runs only on a state change, so an outage produces one "went down"
transition; the cooldown dedup dropped it, and nothing fired again while
the monitor stayed down. On an escalation channel no page opened. The
recovery was still sent. The dedup key is now forgotten when the monitor
recovers, so the cooldown covers one outage, as the docs describe it
("the same ongoing condition"), and the next outage alerts.

A page was left open when its monitor recovered while the rule was
snoozed: the snooze check ran before the recovery handling and skipped the
auto-resolve along with the notifications. Ending an outage is not a
notification, so the auto-resolve and the dedup reset now run before the
snooze check. A snoozed rule still sends nothing.

Closes #387
Closes #391

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@FrameAutomata FrameAutomata added the ci Run CI on this PR (remove and re-add to re-validate after a push) label Oct 4, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci Run CI on this PR (remove and re-add to re-validate after a push)

Projects

None yet

1 participant