Overview
Every controller re-posts each live pod on its node once a minute (POST /pod/spec). A row not re-posted for PEER_STALE_ALIVE_SECS
(default 900) is marked dead by the broker’s stale-alive sweep, and the
pod disappears from /pod/info, /pod/namespaces, the Workloads view
and the maps. When a controller is running but its resync has stopped,
that removal is the only symptom. GET /node/status names the nodes,
so the UI can say “3 nodes have not reported pods since 07:37” instead
of silently showing fewer pods.
GET /node/status
Requires theread scope when broker auth is on. One row per node the
broker has heard from, in node-name order. Derived from tables the
controller already writes; nothing new is posted.
Example
Response
A node with
stale: true and a recent lastHeartbeatAt is a running
controller that has stopped re-posting pods, the case the UI banner
reports; with lastPodPostAt: null as well it is one that has not posted
for longer than retention keeps dead rows, or never could. A node with
both stale is one that left the cluster; its rows age out through
retention.
The sweep itself logs at WARN when it marks rows dead, with the per-node
counts (by_node=ip-10-62-65-125=55 ip-10-62-76-240=55), so the same
event is visible in the broker log.