Skip to main content

What is Seccomp?

Seccomp (Secure Computing Mode) is a Linux kernel feature that restricts which system calls (syscalls) a process can make.

Why Limit Syscalls?

Most applications use only 50-100 of Linux’s 300+ syscalls. Blocking unused syscalls:
  • Reduces attack surface
  • Prevents privilege escalation exploits
  • Stops malicious code from using dangerous syscalls

How kguardian Generates Profiles

  1. Captures syscalls per pod via eBPF, at the cluster’s capture tier
  2. Aggregates unique syscall names per workload over the observation period
  3. Generates a JSON profile with the allowlist

Capture tiers

The controller does not necessarily see every syscall — it traces at one of five tiers, set cluster-wide with syscalls.captureLevel in Helm and raisable per workload with the kguardian.dev/syscall-capture pod annotation: Only a full capture yields a profile that is safe to enforce. A profile is an allowlist: anything not in it is denied, so a profile built from a partial tier denies the syscalls the tier never traced. kguardian marks such profiles as partial capture in the UI, in the exported manifest, and in the SeccompProfile CR’s CaptureComplete condition. The lower tiers exist for monitoring clusters that never intend to enforce; full is the default because syscalls are de-duplicated per pod inside BPF, so its cost is close to low.

Startup coverage

A container makes some of its syscalls only while it starts: the runtime’s setup and the application’s first few hundred milliseconds. The controller captures these from the moment the container’s cgroup is created, before the pod is registered, and adds them to the pod’s set once the pod watcher has registered it (at the pod’s tier). A profile recorded from a fresh pod is therefore startup-complete. Two things are left out on purpose. The container runtime’s rootfs, hostname, keyring and namespace setup (mount, pivot_root, sethostname, keyctl, unshare, …) runs before the runtime installs the container’s seccomp filter, so it never needs to be allowed. What the runtime does after installing the filter (capset, setresuid, execve, …) is kept, because an enforced profile has to allow it. The pod sandbox (pause) container has its own seccomp profile, and its syscalls are not merged into the pod’s set. A pod is registered once it is Ready. A container waits up to STARTUP_CAPTURE_KNOWN_POD_TTL_SECONDS (default 3600) for its pod to become Ready. It waits up to STARTUP_CAPTURE_PENDING_TTL_SECONDS (default 600) if the pod watcher does not know the pod. A controller with startup capture logs Startup syscall capture attached at start. On older controllers, or where that line is replaced by a warning, trigger at least one container restart under SCMP_ACT_LOG before enforcing a profile recorded from fresh pods.

Whose syscalls count

A syscall is credited to a pod only when the calling process is in one of that pod’s cgroups. Host processes that enter a pod’s network namespace are excluded, as are the syscalls they make there. These include containerd pinning and unmounting the namespace (mount, umount2), CNI plugins on pod setup and teardown, and anything else that setns()es in. Before this change they were credited to the pod, so most profiles allowed mount and umount2. See the upgrade note in the chart’s UPGRADING.md. hostNetwork pods share the node’s network namespace and are told apart the same way, by cgroup, so each gets its own set. Static pods are matched through their config hash. A pod deleted and recreated under the same name starts a fresh set; nothing recorded for the previous pod carries over. Known limitations:
  • The pod sandbox (pause) container’s steady-state syscalls are still part of the pod’s set; only its startup is excluded.
  • Startup capture does not cover static pods: their cgroup carries the config hash, and the pending path matches the cgroup’s UID to a registered pod.

From observation to a node

kguardian never applies a profile itself. It exports a SeccompProfile manifest you commit and apply; the controller on every node then writes the file that CR describes and reports readiness and drift in its status. See policy as code and the distribution guide. Example generated profile:

Actions

  • SCMP_ACT_ALLOW: Allow the syscall
  • SCMP_ACT_ERRNO: Block with error (default for unlisted)
  • SCMP_ACT_LOG: Log the syscall but allow it
  • SCMP_ACT_KILL: Kill the calling thread
  • SCMP_ACT_KILL_PROCESS: Kill the whole process (most restrictive)

Denials: The Kernel’s Own Verdict

SCMP_ACT_LOG writes every would-be denial to the node’s kernel audit log (type=SECCOMP in dmesg / auditd) — that is what audit mode has always meant. Until now kguardian never read that back: the audit output of the very profile it generated was invisible to it, so knowing whether a profile was ready to enforce meant an operator going and reading node kernel logs by hand. kguardian closes that loop with a kprobe on the kernel’s own audit_seccomp(), the function every loggable seccomp verdict passes through — unconditionally for SCMP_ACT_LOG and the KILL actions, and for SCMP_ACT_ERRNO only when the filter was loaded with SECCOMP_FILTER_FLAG_LOG, which the controller sets on every enforcing profile it renders so promotion never turns denials invisible. The controller attributes each verdict to a pod and workload, the broker stores it, and it comes back as kguardian_seccomp_denial* Prometheus metrics and the SeccompProfile’s status.denials / DenialsObserved condition. See distributing profiles for how this gates promotion in practice. This is a different question from the Drift condition, and the two are easy to conflate because both are about “a syscall the CR doesn’t cover”: A workload can drift heavily while producing zero denials, simply because nothing has referenced its profile yet. A profile can also show no drift at all — the observed set and spec.syscalls agree completely — and still start producing denials the moment it is promoted to SCMP_ACT_ERRNO, because enforcement (unlike observation) also depends on exact timing, process ancestry, and architecture in ways a syscall-name union can’t capture. Track both; they fail differently. A denial count means nothing without its confidence, so DenialsObserved carries three states rather than two: True means the broker found denials; False with an explicit status.denials.observed: 0 means it looked and found none; Unknown means it answered with nothing to check against — no workloadRef, nothing observed yet, or a response with no denials data at all (predates this feature, or no node is capturing). Treat Unknown as “no data”, never as “clean”. A broker kguardian cannot reach is a sharper edge still: rather than turning Unknown, every condition — this one included — is simply left as it last was, so a CR can read False for hours into an outage and look exactly as clean as before it started. See promote to enforcing for exactly how that distinction gates a real decision. Denial capture does not depend on syscalls.captureLevel — it reads the kernel’s own decision, not traced syscalls — so it works at every capture tier, including low or custom, and even for a workload with no observed syscall set at all.
Next steps: