> ## Documentation Index
> Fetch the complete documentation index at: https://docs.kguardian.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Node catalog

> The nodeCatalog Helm values, the Controller's NODE_CATALOG_* env vars, how it verifies the worker and the container root, its kernel needs and its limits

Reference for node SBOM cataloging: the Controller opens each running
container's root filesystem and hands it to the `cataloger` sidecar, which
returns the image's packages. The feature is off by default; with
`nodeCatalog.enabled=false` the chart renders exactly what it rendered
before the feature existed.

## Enable it

It needs scoped [Broker authentication](/installation) and a `catalog` key
in the same Secret. The key is optional to the pods: without it the Broker
answers the catalog routes `503` and the Controller idles, nothing crash
loops.

```bash theme={null}
kubectl -n kguardian patch secret kguardian-broker-auth \
  -p "{\"stringData\":{\"catalog\":\"$(openssl rand -hex 32)\"}}"

helm upgrade kguardian oci://ghcr.io/kguardian-dev/charts/kguardian -n kguardian \
  --reset-then-reuse-values \
  --set broker.auth.enabled=true \
  --set broker.auth.existingSecret=kguardian-broker-auth \
  --set nodeCatalog.enabled=true
```

Upgrade with `-f <your values file>` or `--reset-then-reuse-values`
(Helm 3.14 and later). Plain `--reuse-values` renders the new chart over the
previous chart's defaults, so none of the `nodeCatalog` defaults below are
in the values; the chart then falls back to the same defaults, built in.
With `--set`, use `--set-string` for a value that looks like a number but
must stay a string (an image tag such as `1.2`, a Secret key name such as
`2026`), or `--set-json` for an object such as `nodeCatalog.worker.appArmorProfile`.

Enabling restarts the Controller DaemonSet and the Broker once (new env
vars and the sidecar). Rotating the `catalog` key needs both restarted
again. It also needs Broker and Controller releases that include the node
catalog (the releases after Broker 1.19.7 and Controller 1.16.2) and the
`cataloger` image; an older Broker answers `404` and the Controller idles.

The chart refuses to render when `nodeCatalog.enabled=true` and
`broker.auth` is off or in `shared` mode, when `nodeCatalog.epoch` is above
`nodeCatalog.maxEpoch`, when `nodeCatalog.maxHoldSeconds` does not outlast
`nodeCatalog.scanTimeoutSeconds`, and when the worker's memory limit cannot
hold `memoryLimit` + `tmpLimit` + 128Mi. `values.schema.json` checks the
types and ranges below.

Check progress with `GET /catalog/coverage` and `GET /catalog/status?node=<name>`
on the Broker (read scope). With `broker.metrics.prometheusRule.enabled`
the chart also ships the `kguardian-broker-node-catalog` PrometheusRule:
token missing, `lsm_denied` skips, scan failures (`timeout`, `oom`,
`error`) above 20% of grants, grants with no completed SBOM in 6 hours
that no node skipped (silenced while `grants=false`), queued work with no
grant for 25 hours, and scans deferred by node pressure for 2 hours.

**Kill switch:** `--set nodeCatalog.grants=false` stops every new grant at
the Broker (a Broker restart, no Controller restart). Uploads under a live
lease still complete.

## Helm values

| Key | Type | Default | Renders |
| - | - | - | - |
| `nodeCatalog.enabled` | bool | `false` | The sidecar, the Controller's `NODE_CATALOG=on` and the env below, `BROKER_TOKEN_CATALOG` on both pods. |
| `nodeCatalog.epoch` | int | `1` | Controller `NODE_CATALOG_EPOCH`. 1 to `maxEpoch`. |
| `nodeCatalog.grants` | bool | `true` | Broker `NODE_CATALOG_GRANTS` (the kill switch). |
| `nodeCatalog.retentionDays` | int | `14` | Broker `NODE_CATALOG_RETENTION_DAYS`; `0` keeps rows. |
| `nodeCatalog.maxEpoch` | int | `1000` | Broker `NODE_CATALOG_MAX_EPOCH`. |
| `nodeCatalog.maxHoldSeconds` | int | `7200` | Broker `NODE_CATALOG_MAX_HOLD_SECS`. At least 900, above `scanTimeoutSeconds`. |
| `nodeCatalog.scanTimeoutSeconds` | int | `600` | Controller `NODE_CATALOG_SCAN_TIMEOUT_SECS`, 10 to 1800. |
| `nodeCatalog.maxFiles` | int | `2000000` | Controller `NODE_CATALOG_MAX_FILES`, at most 10 000 000 (the worker's ceiling). |
| `nodeCatalog.maxComponents` | int | `50000` | Controller `NODE_CATALOG_MAX_COMPONENTS`, at most 50 000. |
| `nodeCatalog.minScanIntervalSeconds` | int | `30` | Controller `NODE_CATALOG_MIN_SCAN_INTERVAL_SECS`, at least 30. |
| `nodeCatalog.pressureThreshold` | int | `40` | Controller `NODE_CATALOG_PRESSURE_THRESHOLD`, 0 to 100. |
| `nodeCatalog.maxPressureDeferSeconds` | int | `1800` | Controller `NODE_CATALOG_MAX_PRESSURE_DEFER_SECS`, at most 86 400. |
| `nodeCatalog.readOnlyClone` | bool | `false` | Controller `NODE_CATALOG_RO_CLONE`. |
| `nodeCatalog.worker.image.*` | | `ghcr.io/kguardian-dev/kguardian/cataloger`, tag `v0.1.0` | The sidecar image; `sha` pins a digest. |
| `nodeCatalog.worker.memoryLimit` | string or int | `640Mi` | Worker `CATALOG_MEMORY_LIMIT`, rendered in bytes. Whole bytes or a whole number of `Ki`, `Mi`, `Gi` or `Ti`. |
| `nodeCatalog.worker.tmpLimit` | string or int | `192Mi` | Worker `CATALOG_TMP_LIMIT` and the `/tmp` emptyDir `sizeLimit`, in bytes. Same forms. |
| `nodeCatalog.worker.resources` | object | requests 50m / 512Mi, limits 500m / 1Gi | Sidecar resources (below). Override, don't null: `null` restores the default. |
| `nodeCatalog.worker.seLinuxOptions` | object | `{type: container_t, level: "s0-s0:c0.c1023"}` | Sidecar SELinux label (below). Override, don't null: `null` restores the default and `{}` removes the label. |
| `nodeCatalog.worker.appArmorProfile` | object | `{}` | Sidecar AppArmor profile (Kubernetes 1.30+). |
| `nodeCatalog.worker.logLevel` | string | `info` | Worker `LOG_LEVEL`. |
| `broker.auth.keys.catalog` | string | `catalog` | Secret key holding the catalog token. |

The Controller's other `NODE_CATALOG_*` variables (below) keep their
defaults; the chart pins `NODE_CATALOG_SOCKET` to
`/run/kguardian/catalog/worker.sock` and sets `POD_UID` from the downward
API.

## The sidecar

The `cataloger` container in the Controller pod runs capability model (i)
from `cataloger/README.md`:

* `runAsUser: 0`, `runAsNonRoot: false`, `allowPrivilegeEscalation: false`,
  `readOnlyRootFilesystem: true`, `seccompProfile: RuntimeDefault`;
  capabilities `drop: [ALL]`, `add: [DAC_READ_SEARCH, SETUID, SETGID]`. The
  uid-0 parent parses nothing; every scan runs in a fresh child as uid
  2000000000 with `DAC_READ_SEARCH` alone, under seccomp.
* Network: the Controller pod is `hostNetwork`, so the sidecar shares the
  node's network namespace. The uid-0 parent opens only its Unix socket,
  but nothing but its own code stops it opening others. The scan child,
  the only process that parses image content, has a seccomp filter that
  denies `socket` and `socketpair` of every family.
* No token: the pod mounts the service-account token for the Controller,
  and the sidecar gets an empty directory over
  `/var/run/secrets/kubernetes.io/serviceaccount` instead. No Broker token,
  no host mounts.
* `/run/kguardian/catalog`: an emptyDir shared with the Controller. The
  worker makes it `0700` and the socket `0600`, both owned by uid 0, and the
  Controller refuses anything else. An emptyDir is created owned by uid 0,
  and `fsGroup` on `controller.podSecurityContext` changes only its group
  and mode, never the owner, so the check holds with or without it (the
  worker's `fchmod` to `0700` also clears the set-group-ID bit kubelet adds).
* `/tmp`: a memory-backed emptyDir with `sizeLimit` equal to `tmpLimit`.
  Its pages count against the sidecar's memory limit, which is why the
  limit must hold `memoryLimit` + `tmpLimit` + 128Mi: the worker's own
  limits then trip before a cgroup OOM kill, which on cgroup v2 takes the
  whole container (kubelet sets `memory.oom.group`). Nodes with kubelet's
  `singleProcessOOMKill: true` (Kubernetes 1.32 and later) kill only the
  scan child.
* Liveness runs `kguardian-cataloger ping`, which the worker answers even
  while a scan runs; there is no readiness probe. The worker exits instead
  of idling when it cannot run safely (a bad `CATALOG_*` setting, a
  capability set it cannot reduce, a uid-0 parent without `SETUID`/`SETGID`,
  a scan child that fails to start, a socket directory not owned by uid 0).
  Kubernetes then restarts it, and while it is not running the Controller
  pod reports NotReady, which also holds a DaemonSet rollout on that node.
  The Controller itself keeps capturing; cataloging on that node reports
  `worker_unavailable`.
* Memory: requests 512Mi, limits 1Gi. The request is reserved on every
  node. Under node memory pressure the kubelet evicts pods using more than
  their request first, and this is the Controller's pod: a lower request
  saves memory per node but makes the pod an earlier eviction candidate
  while a large scan runs; a request equal to the limit avoids that at 1Gi
  per node.

**SELinux.** The default `container_t` with `s0-s0:c0.c1023` lets the
worker read another container's files through the handle it is given: the
category range lifts MCS separation only, and `container_t` stays the
confined container domain. With the default per-pod label every read is
denied (`lsm_denied`). Never use `super_t`, `spc_t` or `control_t`; the
schema refuses them. Nodes without SELinux ignore the setting.

**Cloud credentials.** Workload-identity webhooks inject credentials into
every container of a pod whose service account carries an identity. If the
Controller's service account is given one, the chart keeps it out of the
sidecar with pod annotations the webhooks honour:
`eks.amazonaws.com/skip-containers: cataloger` (the Amazon EKS pod identity
webhook, which serves IRSA) and `azure.workload.identity/skip-containers:
cataloger` (Azure Workload Identity). Check your EKS Pod Identity add-on
version honours the same annotation before relying on it. GKE Workload
Identity injects nothing, but it does not apply to `hostNetwork` pods: they
use the node's service account through the metadata server. The same holds
on every cloud: from the node's network namespace the sidecar's parent can
reach the instance metadata service (EC2, Azure, GCE) and the node's
credentials, exactly as the Controller already can. The scan child cannot
(no sockets). Where that matters, restrict metadata access at the node.

**Pod Security.** The Controller pod is already privileged, so its
namespace must allow the `privileged` Pod Security level; the sidecar asks
for less than the Controller container does. The chart ships no
NetworkPolicy or admission policy that selects the Controller pod.

## Environment variables

**Controller** (`kguardian-controller` DaemonSet)

With `NODE_CATALOG` unset or off, the Controller reads no other variable in
this table and starts nothing: no task, no socket, no Broker call.

| Env | Default | Description |
| - | - | - |
| `NODE_CATALOG` | off | `on` / `true` / `1` turns cataloging on. |
| `BROKER_TOKEN_CATALOG` | unset | The Broker's `catalog` token. Without it the Controller logs once and catalogs nothing. |
| `NODE_CATALOG_SOCKET` | `/run/kguardian/catalog/worker.sock` | The worker's Unix socket, on the `emptyDir` both containers mount. |
| `NODE_CATALOG_EPOCH` | `1` | Catalog generation. Raising it re-catalogs every digest. Must not exceed the Broker's `NODE_CATALOG_MAX_EPOCH`. |
| `NODE_CATALOG_RO_CLONE` | off | Pass the worker a read-only, submount-free `open_tree` clone of the root instead of the plain `O_PATH` handle. Falls back to the handle when the kernel refuses, or when the clone is not the same filesystem object as the verified root. |
| `NODE_CATALOG_HOST_PROC` | `COMPUTE_HOST_PROC`, else `/proc` | Where the host's `/proc` is mounted. |
| `NODE_CATALOG_SCAN_TIMEOUT_SECS` | `600` | Worker scan deadline, 10 to 1800. |
| `NODE_CATALOG_MAX_FILES` | `2000000` | Filesystem entries the worker indexes. |
| `NODE_CATALOG_MAX_COMPONENTS` | `50000` | Components per SBOM (at most 50 000). |
| `NODE_CATALOG_MAX_RESPONSE_BYTES` | `16777216` | Largest worker response accepted, 1 MiB to 64 MiB. It bounds the Controller's memory (below). |
| `NODE_CATALOG_MIN_SCAN_INTERVAL_SECS` | `30` | Minimum gap between claims (at least 30). |
| `NODE_CATALOG_IDLE_SECS` | `600` | An unchanged offer that got no grant is re-sent after this. |
| `NODE_CATALOG_STARTUP_JITTER_SECS` | `120` | The first claim waits a random time up to this (at most 3600), so a fleet restart does not arrive at once. |
| `NODE_CATALOG_PRESSURE_THRESHOLD` | `40` | Host PSI `some avg10` percent (cpu, io or memory) above which scans wait. |
| `NODE_CATALOG_MAX_PRESSURE_DEFER_SECS` | `1800` | Longest a scan waits for pressure to drop (at most 86400). |
| `NODE_CATALOG_PEER_UIDS` | `0` | Comma-separated uids the worker process may run as. |
| `POD_UID` | unset | The Controller pod's own `metadata.uid` (downward API). Its images are never offered. Without it the Controller reads the UID from its kubelet mounts. |

Values over a maximum are clamped with a warning.

## How the worker is verified

Anything in the pod that can write to the shared `emptyDir` could plant a
socket there and receive container root handles, so the Controller pins
the socket to the worker before every connection
(`cataloger/PROTOCOL.md` section 1.1):

1. **The socket file.** The socket is opened `O_PATH|O_NOFOLLOW` and must be a
   socket owned by uid 0 with mode `0600`; its directory is opened
   `O_PATH|O_NOFOLLOW|O_DIRECTORY` and must be owned by uid 0 with mode
   `0700`. Symlinks are never followed. The connect goes through
   `/proc/self/fd/<n>` of that `O_PATH` handle, so the file that was checked
   is the one connected to. This proves the socket was created by a uid-0
   process with the `emptyDir` mounted.
2. **The peer's uid.** After connecting, `SO_PEERCRED` must report an allowed
   uid (`0`).
3. **The peer's container** (Linux 6.5 and later). `SO_PEERPIDFD` pins the
   worker process, and its cgroup must be a sibling of the Controller's own
   container, which means the same pod.

The kernel is probed once at startup. On kernels before 6.5, steps 1 and 2
are the check; nothing is disabled at runtime, and a pidfd error on a
6.5+ kernel rejects that one connection only.

## How the container root is verified

For a granted digest, the Controller picks a running container with it and:

* asks containerd for the task's pid, and for the upperdir of the container's
  snapshot (`Containers.Get`, then `Snapshots.Mounts`);
* opens `/proc/<pid>` as a handle that dies with the process, checks the
  cgroup (pod and container) and start time, opens `root` through it, and
  checks the identity again (pid reuse cannot swap the process);
* requires the `/` entry of the container's mountinfo to be a whole overlay
  mount (root field `/`) with an upperdir, requires the root handle to be
  on that mount (`statx` `STATX_MNT_ID`), and requires the upperdir to be
  the one containerd reports for the container's snapshot. Anything else is
  `unsupported_rootfs`; lazily pulled roots (stargz, SOCI, nydus, FUSE) are
  `lazy_snapshotter`. If containerd cannot answer (unreachable, an error, a
  timeout), that is a fact about the node, not the image: the next replica
  is tried, and if none works the grant ends `pid_gone`.

Only containerd is supported, so every root is checked against its snapshot.

## Drift

A container that changed its packages at runtime is not the image. The
Controller looks at the four package databases (`lib/apk/db/installed`,
`var/lib/dpkg/status`, `var/lib/rpm`, `usr/lib/sysimage/rpm`) in the
container's upperdir, resolved under the host root with `RESOLVE_IN_ROOT`,
one name at a time with no symlinks followed. These are drift, and the
next replica is tried:

* the database written, or deleted (an overlay whiteout);
* a parent directory deleted (a whiteout), replaced by a file or symlink,
  or deleted and recreated (`trusted.overlay.opaque` or
  `user.overlay.opaque` set to `y` or `x`).

If the upperdir cannot be read or an xattr read fails, drift is unknown
and the SBOM is `partial` (`drift_unknown`). The same applies when the
Controller lacks `CAP_SYS_ADMIN` (it is privileged by default): without it
`trusted.overlay.*` reads as absent, so it logs once at startup and every
SBOM from the node is `partial`.

**Language packages.** Deleting a Python, Node or Ruby package does not
touch those databases. The Controller also looks for whiteouts or opaque
directories in the system-wide locations only:

* `usr/lib/python*/{site,dist}-packages` and `usr/local/lib/python*/{site,dist}-packages`;
* `usr/lib/node_modules` and `usr/local/lib/node_modules`;
* `usr/local/bundle/gems`, `var/lib/gems/*/gems` and `usr/local/lib/ruby/gems/*/gems`.

Each directory is read at most 4096 entries deep, and at most 65 536
entries are read in all. A hit makes the SBOM `partial` (`lang_whiteout`);
a directory that could not be read completely makes it `partial` too
(`lang_whiteout_unknown`). Deletions in application-local trees
(`/app/node_modules`, a virtualenv) are **not** detected.

## Kernel requirements

| Need | Kernel | Without it |
| - | - | - |
| `SO_PEERPIDFD`, for the peer cgroup check | Linux 6.5 | The worker is verified by its socket file and uid only. |
| `statx` `STATX_MNT_ID`, to tie the root handle to the `/` mount | Linux 5.8 | `kernel_unsupported` (the node does not catalog). |
| `openat2` with `RESOLVE_IN_ROOT`, for the drift check | Linux 5.6 | Drift is unknown; SBOMs from the node are `partial`. |
| `open_tree` / `mount_setattr`, only for `NODE_CATALOG_RO_CLONE` | Linux 5.12 | Falls back to the plain root handle. |

## Memory

One scan runs at a time. The worker's response is at most
`NODE_CATALOG_MAX_RESPONSE_BYTES` (16 MiB by default). It is parsed into
bounded structures that discard over-long strings, invalid paths and
anything past the caps while parsing, and the raw response is freed as
soon as parsing ends. The upload moves components into one page at a time
(at most 7 MiB encoded). The worst-case extra memory is about the response
size plus 2.5 times it for the parsed package list: roughly 56 MiB at the
default, and roughly 220 MiB at the 64 MiB ceiling. Size the Controller's
memory limit for that before raising it.

## A node the Broker refuses

If the Broker answers a claim with 422 (for example `NODE_CATALOG_EPOCH`
above the Broker's `NODE_CATALOG_MAX_EPOCH`), the Controller logs one
`error` naming both settings and stops claiming. It keeps sending an empty
offer (epoch 0) every `NODE_CATALOG_IDLE_SECS`, so the Broker still records
the node and its platform with no claims, and repeats a `warn` at most once
an hour. Once an hour it sends the empty offer under its real epoch
instead; when the Broker accepts that (its limit was raised), claiming
resumes without a restart. `GET /catalog/status?node=<name>` shows such
a node as seen recently, with no claims.

A 503 from the claim route (catalog not configured on the Broker) backs off
from 60 s, doubling, to at most 30 minutes, and is logged once.
