Cluster watchdog
A CronJob in the cluster-watchdog namespace that runs every 5 minutes and tells a human when the cluster is quietly broken: a node that stays NotReady, or a Longhorn volume that is faulted or stays degraded. It exists because rpi02 died on 2026-08-14 (its NVMe dropped off the bus, the kubelet and API server went with it, but the board kept running from memory and kept answering ARP for every kube-vip virtual IP) and sat like that for four weeks with every volume degraded, Authentik down and media playback flapping, and nothing said a word.
Manifests and the script live next to this page: cluster-watchdog.yaml (namespace, read-only RBAC, CronJob), watchdog.py (the check), and kustomization.yaml (wraps the script in a hash-suffixed ConfigMap so an edit rolls out on the next apply).
What it checks
| Check | Condition | Grace period |
|---|---|---|
| Node | Ready condition is not True |
NODE_GRACE_MINUTES, 10 min, so a reboot does not page you |
| Longhorn volume | robustness: faulted |
none, all replicas are gone |
| Longhorn volume | robustness: degraded while attached |
LONGHORN_GRACE_MINUTES, 60 min, so the replica rebuild after a node returns does not page you |
Detached volumes with robustness unknown are normal and ignored. Every run also prints a one-line cluster summary to its log, so kubectl logs shows what the watchdog saw even when nothing was wrong.
How it alerts
State lives in the ConfigMap cluster-watchdog-state (the script creates it). Each problem gets one entry with first_seen and last_notified, and a run sends at most one message:
- NEW when a problem appears.
- STILL as a reminder every
REMIND_HOURS(24 h) while it persists. - RECOVERED when it clears, with what the problem was.
The message goes through Azure Communication Services (ACS), signed with the same access-key HMAC scheme as the media SMS notifier (media-sms-notifier in the servarr namespace). Email is the primary channel; SMS is optional. Without the acs-notify Secret the job runs in log-only mode: it logs the message it would have sent and still tracks state. When the Secret appears later, the next run notices the channel change and re-sends every open problem once.
Configure the delivery channel
Create the acs-notify Secret in the cluster-watchdog namespace. Every key is optional; email needs the first four, SMS needs the endpoint and key plus the two SMS keys. Values come from the ACS resource in the Azure portal (endpoint and access key under Keys, the sender address under Email > Domains, the number under Phone numbers):
kubectl -n cluster-watchdog create secret generic acs-notify \
--from-literal=ACS_ENDPOINT='https://<resource>.communication.azure.com' \
--from-literal=ACS_ACCESS_KEY='<access key>' \
--from-literal=ACS_EMAIL_FROM='DoNotReply@<domain>.azurecomm.net' \
--from-literal=ACS_EMAIL_TO='you@example.com' \
--from-literal=ACS_SMS_FROM='+4512345678' \
--from-literal=ACS_SMS_TO='+4587654321'
The endpoint, key and SMS number are the same ones the media SMS notifier already uses (acs-sms in the servarr namespace), so SMS alerts can be switched on today by copying those three values; email needs an email domain connected to the ACS resource first.
To change the Secret later, delete and recreate it. The next run picks it up.
Apply and verify
kubectl apply -k docs/k3s/cluster-watchdog/
Run it on demand instead of waiting for the schedule, then read the log:
kubectl -n cluster-watchdog create job --from=cronjob/cluster-watchdog manual-1
kubectl -n cluster-watchdog wait --for=condition=complete --timeout=180s job/manual-1
kubectl -n cluster-watchdog logs job/manual-1
kubectl -n cluster-watchdog delete job manual-1
A healthy cluster logs one line like nodes 3/3 Ready; Longhorn volumes: 15 healthy, 2 unknown; problems=0 new=0 reminders=0 recovered=0. The scheduled runs keep their last three Jobs, so kubectl -n cluster-watchdog logs -l app=cluster-watchdog --tail=5 shows recent history.
Tuning is by environment variable on the CronJob (NODE_GRACE_MINUTES, LONGHORN_GRACE_MINUTES, REMIND_HOURS); edit cluster-watchdog.yaml and re-apply.
Permissions
The ServiceAccount has a cluster-wide read-only grant on nodes and Longhorn volumes, and one namespaced write grant: get/update pinned to the cluster-watchdog-state ConfigMap plus create (which cannot be name-scoped) in its own namespace. It cannot restart, delete or patch anything.
Limits
- It runs inside the cluster, so it cannot report a cluster that is entirely down or an API server that is unreachable. A dead man's switch (an external service that pages when the watchdog stops checking in) would close that gap and is not built.
- It watches nodes and Longhorn only. Stuck
Terminatingpods, crash-looping deployments and failed certificates are symptoms it does not look at. - A message that fails to send is retried on the next run because state only advances after a successful delivery.
- Multi-channel delivery is best-effort per run, not per channel. A run counts as delivered if at least one configured channel succeeds, so if both email and SMS are set and only the SMS send fails, that SMS copy is not independently retried before the next reminder. This avoids re-sending to the working channel every five minutes when a secondary channel is persistently down; email is the primary channel.
- A node that flaps just under the grace period may not alert. NotReady detection uses the node's own
lastTransitionTime, which resets on each flap, so a node oscillating Ready/NotReady faster thanNODE_GRACE_MINUTEScan stay unhealthy without ever tripping the alert. A steadily-down node (the common failure, e.g. rpi02) is caught normally. - A degraded volume is judged by Longhorn's
lastDegradedAt. If that field is absent the volume is not alerted, so a rebuild that has only just started does not page during its normal grace window. Longhorn sets the field reliably (the rpi02 volumes had it), so a genuinely long degradation is still caught; a volume permanently missing the field would be the rare gap. Afaultedvolume has no such dependency and alerts immediately. - Email delivery is confirmed only as far as Azure accepting it. The ACS email API returns 202 Accepted (queued); a later asynchronous failure (bounce, suppressed sender) is not detected, since that would require polling the operation status. SMS is synchronous. Watch for an unexpected quiet spell as the signal that a channel has silently stopped.
- The ACS request signer is duplicated with the media SMS notifier (
servarrnamespace) rather than shared, since the two run as independent workloads in different namespaces; an ACS signing-scheme or API-version change must be applied in both.