Skip to main content
Reference for node SBOM cataloging: the Controller opens each running container’s root filesystem and hands it to the cataloger sidecar, which returns the image’s packages. The feature is off by default. The Helm values that render these variables arrive with the chart change; until then, set them on the Controller DaemonSet directly.

Environment variables

Controller (kguardian-controller DaemonSet) With NODE_CATALOG unset or off, the Controller reads no other variable in this table and starts nothing: no task, no socket, no Broker call. Values over a maximum are clamped with a warning.

How the worker is verified

Anything in the pod that can write to the shared emptyDir could plant a socket there and receive container root handles, so the Controller pins the socket to the worker before every connection (cataloger/PROTOCOL.md section 1.1):
  1. The socket file. The socket is opened O_PATH|O_NOFOLLOW and must be a socket owned by uid 0 with mode 0600; its directory is opened O_PATH|O_NOFOLLOW|O_DIRECTORY and must be owned by uid 0 with mode 0700. Symlinks are never followed. The connect goes through /proc/self/fd/<n> of that O_PATH handle, so the file that was checked is the one connected to. This proves the socket was created by a uid-0 process with the emptyDir mounted.
  2. The peer’s uid. After connecting, SO_PEERCRED must report an allowed uid (0).
  3. The peer’s container (Linux 6.5 and later). SO_PEERPIDFD pins the worker process, and its cgroup must be a sibling of the Controller’s own container, which means the same pod.
The kernel is probed once at startup. On kernels before 6.5, steps 1 and 2 are the check; nothing is disabled at runtime, and a pidfd error on a 6.5+ kernel rejects that one connection only.

How the container root is verified

For a granted digest, the Controller picks a running container with it and:
  • asks containerd for the task’s pid, and for the upperdir of the container’s snapshot (Containers.Get, then Snapshots.Mounts);
  • opens /proc/<pid> as a handle that dies with the process, checks the cgroup (pod and container) and start time, opens root through it, and checks the identity again (pid reuse cannot swap the process);
  • requires the / entry of the container’s mountinfo to be a whole overlay mount (root field /) with an upperdir, requires the root handle to be on that mount (statx STATX_MNT_ID), and requires the upperdir to be the one containerd reports for the container’s snapshot. Anything else is unsupported_rootfs; lazily pulled roots (stargz, SOCI, nydus, FUSE) are lazy_snapshotter. If containerd cannot answer (unreachable, an error, a timeout), that is a fact about the node, not the image: the next replica is tried, and if none works the grant ends pid_gone.
Only containerd is supported, so every root is checked against its snapshot.

Drift

A container that changed its packages at runtime is not the image. The Controller looks at the four package databases (lib/apk/db/installed, var/lib/dpkg/status, var/lib/rpm, usr/lib/sysimage/rpm) in the container’s upperdir, resolved under the host root with RESOLVE_IN_ROOT, one name at a time with no symlinks followed. These are drift, and the next replica is tried:
  • the database written, or deleted (an overlay whiteout);
  • a parent directory deleted (a whiteout), replaced by a file or symlink, or deleted and recreated (trusted.overlay.opaque or user.overlay.opaque set to y or x).
If the upperdir cannot be read or an xattr read fails, drift is unknown and the SBOM is partial (drift_unknown). The same applies when the Controller lacks CAP_SYS_ADMIN (it is privileged by default): without it trusted.overlay.* reads as absent, so it logs once at startup and every SBOM from the node is partial. Language packages. Deleting a Python, Node or Ruby package does not touch those databases. The Controller also looks for whiteouts or opaque directories in the system-wide locations only:
  • usr/lib/python*/{site,dist}-packages and usr/local/lib/python*/{site,dist}-packages;
  • usr/lib/node_modules and usr/local/lib/node_modules;
  • usr/local/bundle/gems, var/lib/gems/*/gems and usr/local/lib/ruby/gems/*/gems.
Each directory is read at most 4096 entries deep, and at most 65 536 entries are read in all. A hit makes the SBOM partial (lang_whiteout); a directory that could not be read completely makes it partial too (lang_whiteout_unknown). Deletions in application-local trees (/app/node_modules, a virtualenv) are not detected.

Kernel requirements

Memory

One scan runs at a time. The worker’s response is at most NODE_CATALOG_MAX_RESPONSE_BYTES (16 MiB by default). It is parsed into bounded structures that discard over-long strings, invalid paths and anything past the caps while parsing, and the raw response is freed as soon as parsing ends. The upload moves components into one page at a time (at most 7 MiB encoded). The worst-case extra memory is about the response size plus 2.5 times it for the parsed package list: roughly 56 MiB at the default, and roughly 220 MiB at the 64 MiB ceiling. Size the Controller’s memory limit for that before raising it.

A node the Broker refuses

If the Broker answers a claim with 422 (for example NODE_CATALOG_EPOCH above the Broker’s NODE_CATALOG_MAX_EPOCH), the Controller logs one error naming both settings and stops claiming. It keeps sending an empty offer (epoch 0) every NODE_CATALOG_IDLE_SECS, so the Broker still records the node and its platform with no claims, and repeats a warn at most once an hour. Once an hour it sends the empty offer under its real epoch instead; when the Broker accepts that (its limit was raised), claiming resumes without a restart. GET /catalog/status?node=<name> shows such a node as seen recently, with no claims. A 503 from the claim route (catalog not configured on the Broker) backs off from 60 s, doubling, to at most 30 minutes, and is logged once.