Monitoring and alerts for BYOC databases
Nothing in this CLI pushes. Every observability surface it carries is a read you make, so monitoring is built by polling those reads on a schedule, turning each answer into a verdict, and letting an exit status carry the verdict out to whatever runs the schedule. The logs and metrics guide documents the reading itself, and this page covers what to build on top of it.
What the CLI exposes
An alert rule draws on metrics, logs, a resource's own record, the task list and the connection diagnostic, and every one of them is a read that prompts for nothing and writes nothing. The following table describes each command:
| Command | What it reads |
|---|---|
pgedge starfleet byoc database metrics |
The metrics series for one BYOC database, over a lookback set by --interval and narrowed by --columns and --node-name. |
pgedge starfleet byoc database logs |
Log lines from one component on named nodes, with --component-name and --nodes both required. |
pgedge starfleet byoc database get |
The database record, whose status field carries the lifecycle value. |
pgedge starfleet byoc task list |
The tasks behind asynchronous writes, filtered by --subject-id, --subject-kind, --name and --status. |
pgedge starfleet doctor |
The Starfleet connection, at exit 0 whatever it finds, so a rule reads its fields under -o json rather than its status. |
BYOC takes --interval, in value,unit form with second, minute,
hour, day, week, month or year as the unit, singular or plural, and a
comma or a space between the two. The command refuses a zero
value at exit 2 before any request leaves. BYOC's database metrics
command carries no time flags at all.
A metrics poll with a threshold
A threshold rule reads the metrics as JSON, picks one metric out of
the series, and compares it against a number you chose. Under -o
json and -o yaml BYOC answers with the API's own object, a
container of series where each series carries a name, its column names
and its rows, so a filter selects a metric by finding its name among
the columns and reading that position out of a row.
Row order is not published, so select the newest sample by the time
column rather than by taking the last row.
The following script reads one BYOC database's metrics, applies a threshold to the newest sample, and reports a breach on its own exit status:
#!/bin/sh
set -u
db_id="$1"
limit="$2"
pgedge starfleet byoc database metrics "$db_id" \
--interval 15,minutes --profile monitor --timeout 30s \
-o json > metrics.json
rc=$?
if [ "$rc" -ne 0 ]
then
echo "metrics read exited $rc" >&2
exit "$rc"
fi
jq -e --argjson limit "$limit" '
(.series[0] // {}) as $s
| ($s.columns // []) as $c
| ($c | index("time")) as $t
| ($c | index("<metric>")) as $m
| if $t == null or $m == null
or (($s.values // []) | length) == 0
then empty
else ($s.values | max_by(.[$t]))[$m] < $limit
end
' metrics.json > /dev/null
verdict=$?
case "$verdict" in
0)
exit 0
;;
1)
echo "$db_id is over the threshold" >&2
exit 1
;;
*)
echo "$db_id returned no usable sample" >&2
exit 0
;;
esac
The read goes to a file and its status is captured before anything
parses it, so a failed read exits with the CLI's own code before the
filter ever runs. Writing the read as pgedge ... | jq loses that,
because the pipeline reports the filter's status and a failed read
then reaches the rule as an empty input instead of a failure.
The filter's three outcomes are what the case statement branches on:
jq -eexits 0 when the comparison is true, which is the sample sitting under the threshold.jq -eexits 1 when the comparison is false, which is the breach worth alerting on.jq -eexits 4 when the filter produced no output at all, which is an empty window or a metric the response omitted, because a metric with no value anywhere in the window is left out of the response rather than sent as null.
When the lookback comes back empty, BYOC reports a bare sentence on stderr at exit 0.
The <metric> in the filter stands for the column you are
thresholding. Read the series once with -o json and pick the column
name from it. Keep the alert action itself out of the comparison. The
script above says what happened on stderr and exits non-zero, which is
the signal cron, a CI runner or a pager sidecar already knows how to
read.
Alerting on failed operations
Most writes are asynchronous, and without a wait flag a create that later fails to provision still exits 0. The write's own status is therefore not the whole story, and a scheduled sweep of the task list catches the failures a script accepted and walked away from.
task list filters server-side on --status, which takes queued,
running, succeeded or failed. The following script fails when any task
against one database has failed:
#!/bin/sh
set -u
db_id="$1"
pgedge starfleet byoc task list --subject-id "$db_id" \
--status failed --profile monitor --timeout 30s \
-o json > tasks.json
rc=$?
if [ "$rc" -ne 0 ]
then
echo "task list exited $rc" >&2
exit "$rc"
fi
if ! jq -e 'length == 0' tasks.json > /dev/null
then
jq -r '.[] | "\(.name) \(.id) failed"' tasks.json >&2
exit 1
fi
The database record is the second signal, and it answers a different
question: task list says an operation ended badly, while database
get says what state the database was left in. The following script
reads that state and decides on it:
#!/bin/sh
set -u
db_id="$1"
pgedge starfleet byoc database get "$db_id" \
--profile monitor --timeout 30s -o json > db.json
rc=$?
if [ "$rc" -ne 0 ]
then
echo "database get exited $rc" >&2
exit "$rc"
fi
db_status=$(jq -r .status db.json)
if [ -z "$db_status" ] || [ "$db_status" = "null" ]; then
echo "db.json carries no status" >&2
exit 1
fi
case "$db_status" in
available)
exit 0
;;
failed|degraded)
echo "$db_id is $db_status" >&2
exit 1
;;
*)
echo "$db_id is $db_status, not alerting" >&2
exit 0
;;
esac
A BYOC database's status reports one of seven values, and they fall into three groups an alert rule treats differently:
availableis the ready value, and a readiness check compares against it rather than against "notcreating", because a database can reachfailedordegradedwithout passing throughcreatingagain.failedanddegradedare the two worth paging on.queued,creating,modifyinganddeletingdescribe work in flight.
The field is a bare string in the contract, so those seven are the vocabulary the platform uses rather than a fixed set, and a value outside the list is possible. A rule that treats every unrecognized value as a failure pages on one, which is why the script above lets its default branch pass.
A service deployed on a database reports its own lifecycle in state
rather than in status, on a shorter vocabulary, and a script reading
one field where the other lives finds nothing. state is not a
readiness signal: it reads running as soon as the deploy completes,
while the server itself may still be refusing requests, so its one
useful value for an alert is failed. The
BYOC services guide covers what the platform does with
state.
Read the exit status of every one of these calls before reading its output. A plan-entitlement rejection lands on exit 5 alongside a rejected credential, so a rule that retries authentication failures spins on a refusal no credential will ever fix.
Running one of these scripts unattended is covered by scheduling a poll.
Next steps
- The logs and metrics guide documents the metrics and log commands themselves, and the cluster reads that sit underneath them.
- The CI and automation guide covers unattended credentials, complete pipelines on two runners, and scripting against output and exit codes.
- The exit codes guide carries the contract an alert rule branches on and the commands that deliberately depart from it.
- The health checks guide covers the diagnostics to run when a poll fails for a connection reason rather than an operational one.