Skip to content

Alerts

The alert system monitors PostgreSQL metrics and notifies users when thresholds are exceeded or anomalies are detected. This guide explains how alerts appear in the web interface and how to manage the alert lifecycle.

Viewing Alerts

The status panel displays active alerts for the scope selected in the cluster navigator. Alerts appear grouped by severity and category. Each alert shows the rule name, current metric value, threshold, and the server where the alert originated.

Selecting a different node in the cluster navigator updates the alert list to reflect that scope. Estate-wide alerts appear when no specific node is selected.

Severity Levels

Alerts use three severity levels to indicate urgency:

  • A critical alert indicates a severe issue that requires immediate attention.
  • A warning alert indicates a potential problem that should be investigated soon.
  • An info alert indicates a noteworthy condition for general awareness.

The severity level determines the visual styling of each alert in the status panel. Critical alerts appear with red indicators; warning alerts appear with amber indicators; info alerts appear with blue indicators.

Alert Lifecycle

Each alert progresses through a defined set of states from creation to resolution.

Active

The alerter creates an alert with active status when a metric violates a threshold. The alert remains active until the condition resolves or an operator takes action. The system updates the metric value on each evaluation cycle while the alert stays active.

Acknowledged

An operator can acknowledge an active alert to indicate that the issue is under investigation. Acknowledged alerts remain visible in the status panel but move to a separate section. Acknowledging an alert does not resolve the underlying condition.

Cleared

The alerter automatically clears an alert when the triggering condition returns to normal. The alert cleaner runs every 30 seconds and re-evaluates active alerts. When a metric value no longer violates the threshold, the system marks the alert as cleared and records the timestamp.

A gap in the collected data does not clear an alert. When a metric that normally reports a value for every monitored connection stops returning one, the alerter treats the gap as missing data rather than recovery, so the alert stays active and the metric staleness rule reports the probe that stopped. Alerts on metrics that report only whilst the condition holds, such as an inactive replication slot or a blocked session, still clear as soon as the metric goes quiet.

A metric staleness alert that has already fired follows the same principle. When the probe behind such an alert becomes unavailable, because a monitored server no longer offers the extension the probe needs or because the probe run timed out, the alert stays active and its description changes to report that collection has stopped and why. The alert title does not change, so the alert remains recognisable in notification history, and the original description returns once the probe collects again. When an operator disables the probe, or stops monitoring the connection, the alerter clears the staleness alert, because both actions are deliberate, and it restores the original description as the alert clears so that the cleared alert does not report that collection has stopped.

This behaviour holds an existing alert open; it does not raise one. The probe that has stopped is reported by a separate rule, probe_unavailable, which raises a warning alert as soon as a probe that had been collecting stops being available. That moment normally arrives long before a staleness alert could fire, so the two rules cover different halves of the same failure. A probe whose extension was never installed does not raise the alert, which keeps a permanent alert off every probe that has simply never run.

The alerter clears a probe unavailable alert when the probe collects again. The alerter also clears the alert when an operator retires the probe by disabling it or by no longer monitoring the connection, and when the probe no longer reports its availability for that connection.

A held staleness alert and a probe unavailable alert can be open on the same probe at the same time; the first reports that the data went stale and the second reports why, and each one clears on its own terms.

False Positive

An operator can mark an alert as a false positive to indicate that the alert does not represent a real issue. The false positive designation helps refine alert accuracy over time.

Acknowledging Alerts

Click the acknowledge button on an active alert to mark the alert as acknowledged. The alert moves to the acknowledged section of the status panel.

Acknowledging an alert signals to other operators that someone is investigating the issue. The system records the operator who acknowledged the alert and the timestamp.

Alert History

The alert history provides a record of all past alerts for a given scope. Use the alert history to review patterns, identify recurring issues, and verify that resolved conditions remain stable.

Historical alerts include the trigger time, resolution time, peak metric value, and the action taken by the operator.

AI-Powered Analysis

Each alert in the status panel displays a brain icon that triggers an AI-powered analysis. The analysis examines the alert context, historical patterns, and server configuration to produce actionable remediation guidance. See AI Alert Analysis for details on this feature.

Editing Alert Thresholds

Users can edit alert thresholds directly from an alert instance. The edit button on an alert opens the Edit Override dialog for the associated rule and scope. The dialog allows adjustments to the threshold, operator, severity, and enabled state.

The scope dropdown displays the available override levels: server, cluster, and group. The dialog pre-selects the scope that matches the originating context of the alert.

Blackout Interaction

During an active blackout period, the alerter suppresses new alerts for the affected connection or database. Existing active alerts remain visible during a blackout; the blackout only prevents new alerts from being created. See Blackouts for details on maintenance windows.