Prerequisites
Before installing kguardian, verify your environment meets these requirements:Kubernetes Cluster
Kubernetes Cluster
Required:
- Kubernetes version: 1.19+
- Node OS: Linux
- Kernel version: 6.2+ (for eBPF support)
- Container runtime: containerd (required — Docker and CRI-O are not supported)
Helm (Recommended)
Helm (Recommended)
- Helm 3.0+
kubectl Access
kubectl Access
kubectl1.19+ configured with admin privileges
Installation Methods
- Helm (Recommended)
- Kind (Local Development)
Install from OCI Registry
The simplest and recommended method for production deployments.All pods — controller (one per node), broker, db, frontend, and evaluator — should be in
Running state within 2-3 minutes.Install Specific Version
You can pin a particular chart version:Review Default Values
To customize your installation, you can view and save the default values:Configuration Options
Common Helm Values
Customize your installation by creating avalues.yaml file:
charts/kguardian/values.yaml or the auto-generated chart README.
Run More Than One Broker Replica
broker.replicaCount above 1 is safe: the Broker replicas elect a leader
through a coordination.k8s.io/v1 Lease, and only the leader runs the
background jobs that prune or derive shared data (retention, the compute
downsample, the workload profile snapshotter, the stale-pod sweep, the
supply-chain rollups). Every replica serves the API.
leaseDurationSeconds before a follower takes over; on a graceful shutdown
the handover takes a couple of seconds. With leaderElection.enabled: false
every replica runs every job, which duplicates work, so the chart refuses to
render it with replicaCount above 1. The timings must satisfy
leaseDurationSeconds > renewDeadlineSeconds > 1.2 x retryPeriodSeconds,
or the chart refuses to render. See broker_leader and broker_leader_election_active on
/metrics for which replica is leading.
Each replica keeps 4 database connections open and grows its pool up to
broker.dbPoolMaxSize (default 32) under load. During a rolling update the
old and new pods overlap, so the database must accept:
dbPoolMaxSize is raised to broker.audit.inflightPermits + 8 if it is
lower.) With the bundled database the chart does this for you: it leaves
PostgreSQL’s default max_connections of 100 while that is enough and raises
it otherwise (150 for three replicas), and refuses to render an explicit
database.maxConnections below the requirement. With an external database,
check its max_connections against the formula before raising replicaCount;
helm install prints the number in its notes.
Use External PostgreSQL
For production-grade deployments — managed services like AWS RDS / Cloud SQL / a CloudNativePG cluster you already run — disable the bundled Postgres and point the broker at the external instance:AI Assistant
The assistant is optional and off by default. It is one workload — the LLM Bridge — that runs the 23 tools and the policy/seccomp generation in-process against the Broker. Create a Secret holding your provider API key under the keyapi-key, then turn the assistant on with the provider name and the Secret:
OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEY, or GITHUB_TOKEN. For a multi-provider setup, configure each one under llmBridge.secrets.* instead and leave ai.provider empty.
Check it came up with the LLM Bridge health endpoint — hasProvider must be true:
Use an OpenAI-Compatible Gateway
To route the assistant through LiteLLM, vLLM, or an enterprise proxy instead of the vendor API, setai.baseUrl — and, because a gateway’s model ids are its own, usually ai.model too:
ai.provider names. Setting them without ai.provider fails the install rather than being silently ignored.
The variable each value maps to depends on the provider:
Gemini’s asymmetry is deliberate — the key is
GOOGLE_API_KEY while the base URL and model are GEMINI_*, matching each vendor’s own convention. ANTHROPIC_BASE_URL is not new: the Anthropic SDK already honoured it, and the chart now sets it explicitly so a typo is reported against the value rather than surfacing as a connection error.
Errors are diagnostic rather than silent. A malformed or scheme-less base URL fails before any request is sent, naming the variable. A 404 or 405 from an OpenAI-compatible, Copilot, or Gemini gateway names the exact URL that was called and the variable to fix. A whitespace-only value counts as unset, consistently with the API keys.
Serve the MCP Endpoint
The LLM Bridge can serve its 23 tools to external MCP clients atPOST /mcp on its existing port, so Claude Code or another MCP client can query your cluster’s observed telemetry and generate policies from it. It is off by default:
ai.mcp.auth.allowUnauthenticated=true deliberately, because the endpoint serves cluster telemetry to anything that can reach the ClusterIP Service. ai.mcp.rateLimitPerMin moves the per-replica ceiling from its default of 300.
The full client-side walkthrough — port-forward, claude mcp add, .mcp.json, and what the tools answer — is in Connect an MCP Client.
Gate the UI Behind SSO
The UI and the Broker API are unauthenticated by default. In a Gateway API environment running an identity-aware proxy (oauth2-proxy in front of an OIDC provider such as Dex, Keycloak or Authentik), the chart can put the frontend route behind SSO:enabled is false: a SecurityPolicy that ext-auths the frontend route against oauth2-proxy, and an /oauth2/* HTTPRoute so the OIDC callback and /oauth2/userinfo are served same-origin. The UI detects the session from /oauth2/userinfo at runtime, so there is no rebuild and no build flag; without the proxy it stays in local mode.
If SSO is on but the UI still shows local mode, check where /oauth2/userinfo is answered: curl -sI https://<ui-host>/oauth2/userinfo. A 204 with X-Kguardian-Sso: none came from the UI’s own server, so the /oauth2/* route is missing or not attached and the request never reached oauth2-proxy. oauth2-proxy itself answers 200 with a session or 401 without one.
This gates the UI route only. The Broker API is a separate Service. To authenticate it too, turn on scoped broker tokens (broker.auth.enabled). The frontend’s /api proxy then adds the read token on the server side and drops the SSO Authorization header before the request reaches the broker. SSO also doesn’t cover the MCP endpoint above, which has its own bearer token.
SSO guards the Gateway route, not the frontend Service itself. With broker auth on, anything in the cluster that can reach the frontend Service directly can read everything the broker holds, because the proxy adds the read token for it. Restrict ingress to the frontend pods with a NetworkPolicy that admits only your Gateway.
Restrict the UI’s Host Names
The UI server answers only the host names the chart knows:frontend.ingress.hosts (and the TLS hosts) when the chart’s ingress is on, frontend.sso.hostnames when SSO is on, the frontend Service’s in-cluster DNS names (the full one ends in global.clusterDomain, default cluster.local), and anything in frontend.allowedHosts. Any other Host header gets a 403, so a web page cannot rebind its own name to the UI and use the /api proxy, and the read token it adds, from a user’s browser. localhost and IP addresses are always answered, so kubectl port-forward keeps working.
frontend.allowedHosts also accepts one comma-separated string. A wildcard is broader than in Kubernetes: *.corp.example.com matches corp.example.com itself and subdomains at any depth, not just one label.
With none of these set, or with an ingress rule whose host is empty (a catch-all) and no frontend.allowedHosts, the server answers every host, as earlier releases did, and logs a warning at startup.
Install the CLI Plugin
After deploying the controller, install the CLI to generate policies.- Quick Install Script (Recommended)
- Manual Download
- From Source
- Detects your OS and architecture
- Downloads the latest CLI release
- Installs to
/usr/local/bin/kubectl-kguardian - Validates the installation
Verification
After installation, verify all components are working:1. Check Pod Status
2. Check Controller Logs
3. Check Broker Connectivity
4. Test CLI Connection
If all checks pass, kguardian is ready to use!
Upgrade
To upgrade kguardian to a newer version:Upgrading the Broker past 1.19.4 with a bloated Or scale the old Broker to zero for the moment the new one migrates. Either way, the migration then finds an empty table.
pod_compute_latest. The first Broker after 1.19.4 truncates pod_compute_latest in a migration that gives up on its table lock after 5 seconds. If the table is already tens of GB, the old Broker’s queries on it can outlast that on every attempt, and the new Broker pod crash-loops while the old one keeps serving. Before upgrading, run the truncate yourself from psql, retrying until it succeeds (the table is a live cache and refills within one sample interval):Upgrading from an older Broker with large tables. A few migrations add indexes to tables that grow large:
pod_traffic and pod_syscalls (2026-06-01, 2026-07-06, 2026-09-26) and image_sbom_components (2026-09-28). When the table is under 256 MiB the migration builds the index itself. When it is larger the migration skips the build, because a plain CREATE INDEX blocks every insert into the table for as long as it runs and could outlast the new pod’s liveness probe. The leader Broker then builds the same index with CREATE INDEX CONCURRENTLY in the background, which blocks no writes; the log shows building index CONCURRENTLY and index built. Until it finishes, reads that use the index are as slow as they were before the upgrade. The compute history indexes are built the same way, on every install.Uninstall
To completely remove kguardian:This will not remove generated policies that were already applied to your namespaces. Clean those up separately if needed.
Quick Start
Jump right into generating policies
Connect an MCP Client
Point Claude Code or another MCP client at your cluster
Troubleshooting
Fix common installation issues
Architecture
Understand how components work