What is Seccomp?
Seccomp (Secure Computing Mode) is a Linux kernel feature that restricts which system calls (syscalls) a process can make.Why Limit Syscalls?
Most applications use only 50-100 of Linux’s 300+ syscalls. Blocking unused syscalls:- Reduces attack surface
- Prevents privilege escalation exploits
- Stops malicious code from using dangerous syscalls
How kguardian Generates Profiles
- Captures syscalls per pod via eBPF, at the cluster’s capture tier
- Aggregates unique syscall names per workload over the observation period
- Generates a JSON profile with the allowlist
Capture tiers
The controller does not necessarily see every syscall — it traces at one of five tiers, set cluster-wide withsyscalls.captureLevel in Helm and
raisable per workload with the kguardian.dev/syscall-capture pod
annotation:
Only a
full capture yields a profile that is safe to enforce. A profile
is an allowlist: anything not in it is denied, so a profile built from a
partial tier denies the syscalls the tier never traced. kguardian marks such
profiles as partial capture in the UI, in the exported manifest, and in the
SeccompProfile CR’s CaptureComplete condition. The lower tiers exist for
monitoring clusters that never intend to enforce; full is the default
because syscalls are de-duplicated per pod inside BPF, so its cost is close to
low.
Startup coverage
A container makes some of its syscalls only while it starts: the runtime’s setup and the application’s first few hundred milliseconds. The controller captures these from the moment the container’s cgroup is created, before the pod is registered, and adds them to the pod’s set once the pod watcher has registered it (at the pod’s tier). A profile recorded from a fresh pod is therefore startup-complete. Two things are left out on purpose. The container runtime’s rootfs, hostname, keyring and namespace setup (mount, pivot_root, sethostname, keyctl,
unshare, …) runs before the runtime installs the container’s seccomp filter,
so it never needs to be allowed. What the runtime does after installing the
filter (capset, setresuid, execve, …) is kept, because an enforced
profile has to allow it. The pod sandbox (pause) container has its own seccomp
profile, and its syscalls are not merged into the pod’s set.
A pod is registered once it is Ready. A container waits up to
STARTUP_CAPTURE_KNOWN_POD_TTL_SECONDS (default 3600) for its pod to become
Ready. It waits up to STARTUP_CAPTURE_PENDING_TTL_SECONDS (default 600) if
the pod watcher does not know the pod.
A controller with startup capture logs Startup syscall capture attached at
start. On older controllers, or where that line is replaced by a warning,
trigger at least one container restart under SCMP_ACT_LOG before enforcing
a profile recorded from fresh pods.
Whose syscalls count
A syscall is credited to a pod only when the calling process is in one of that pod’s cgroups. Host processes that enter a pod’s network namespace are excluded, as are the syscalls they make there. These include containerd pinning and unmounting the namespace (mount, umount2), CNI plugins on pod
setup and teardown, and anything else that setns()es in. Before this change
they were credited to the pod, so most profiles allowed mount and umount2.
See the upgrade note in the chart’s UPGRADING.md.
hostNetwork pods share the node’s network namespace and are told apart the
same way, by cgroup, so each gets its own set. Static pods are matched through
their config hash.
A pod deleted and recreated under the same name starts a fresh set; nothing
recorded for the previous pod carries over.
Known limitations:
- The pod sandbox (pause) container’s steady-state syscalls are still part of the pod’s set; only its startup is excluded.
- Startup capture does not cover static pods: their cgroup carries the config hash, and the pending path matches the cgroup’s UID to a registered pod.
From observation to a node
kguardian never applies a profile itself. It exports aSeccompProfile
manifest you commit and apply; the controller on every node then writes the
file that CR describes and reports readiness and drift in its status. See
policy as code and the
distribution guide.
Example generated profile:
Actions
SCMP_ACT_ALLOW: Allow the syscallSCMP_ACT_ERRNO: Block with error (default for unlisted)SCMP_ACT_LOG: Log the syscall but allow itSCMP_ACT_KILL: Kill the calling threadSCMP_ACT_KILL_PROCESS: Kill the whole process (most restrictive)
Denials: The Kernel’s Own Verdict
SCMP_ACT_LOG writes every would-be denial to the node’s kernel audit log
(type=SECCOMP in dmesg / auditd) — that is what audit mode has always
meant. Until now kguardian never read that back: the audit output of the
very profile it generated was invisible to it, so knowing whether a profile
was ready to enforce meant an operator going and reading node kernel logs by
hand.
kguardian closes that loop with a kprobe on the kernel’s own
audit_seccomp(), the function every loggable seccomp verdict passes
through — unconditionally for SCMP_ACT_LOG and the KILL actions, and
for SCMP_ACT_ERRNO only when the filter was loaded with
SECCOMP_FILTER_FLAG_LOG, which the controller sets on every enforcing
profile it renders so promotion never turns denials invisible. The controller
attributes each verdict to a pod and workload, the broker stores it, and it
comes back as kguardian_seccomp_denial* Prometheus metrics and the
SeccompProfile’s status.denials / DenialsObserved condition. See
distributing profiles for how this
gates promotion in practice.
This is a different question from the Drift condition, and the two are
easy to conflate because both are about “a syscall the CR doesn’t cover”:
A workload can drift heavily while producing zero denials, simply because
nothing has referenced its profile yet. A profile can also show no drift at
all — the observed set and
spec.syscalls agree completely — and still
start producing denials the moment it is promoted to SCMP_ACT_ERRNO,
because enforcement (unlike observation) also depends on exact timing,
process ancestry, and architecture in ways a syscall-name union can’t
capture. Track both; they fail differently.
A denial count means nothing without its confidence, so DenialsObserved
carries three states rather than two: True means the broker found
denials; False with an explicit status.denials.observed: 0 means it
looked and found none; Unknown means it answered with nothing to check
against — no workloadRef, nothing observed yet, or a response with no
denials data at all (predates this feature, or no node is capturing).
Treat Unknown as “no data”, never as “clean”. A broker kguardian cannot
reach is a sharper edge still: rather than turning Unknown, every
condition — this one included — is simply left as it last was, so a CR can
read False for hours into an outage and look exactly as clean as before
it started. See
promote to enforcing
for exactly how that distinction gates a real decision.
Denial capture does not depend on syscalls.captureLevel — it reads the
kernel’s own decision, not traced syscalls — so it works at every capture
tier, including low or custom, and even for a workload with no observed
syscall set at all.
Next steps: