Skip to main content
This guide covers all installation methods for kguardian, from quick setups to production deployments.

Prerequisites

Before installing kguardian, verify your environment meets these requirements:
Required:
  • Kubernetes version: 1.19+
  • Node OS: Linux
  • Kernel version: 6.2+ (for eBPF support)
  • Container runtime: containerd (required — Docker and CRI-O are not supported)
Verify kernel version:
  • kubectl 1.19+ configured with admin privileges
Verify access:
Kernel Compatibility: kguardian requires Linux kernel 6.2+ for eBPF CO-RE (Compile Once, Run Everywhere) support. Older kernels may work with manual BPF program compilation but are not officially supported.

Installation Methods


Configuration Options

Common Helm Values

Customize your installation by creating a values.yaml file:
Install with custom values:
For the full list of values, see charts/kguardian/values.yaml or the auto-generated chart README.

Run More Than One Broker Replica

broker.replicaCount above 1 is safe: the Broker replicas elect a leader through a coordination.k8s.io/v1 Lease, and only the leader runs the background jobs that prune or derive shared data (retention, the compute downsample, the workload profile snapshotter, the stale-pod sweep, the supply-chain rollups). Every replica serves the API.
Leader election adds a Role in the release namespace that can create Leases and get/update only the Broker’s own, and mounts a service account token into the Broker pod. When the leader dies, the jobs pause for up to leaseDurationSeconds before a follower takes over; on a graceful shutdown the handover takes a couple of seconds. With leaderElection.enabled: false every replica runs every job, which duplicates work, so the chart refuses to render it with replicaCount above 1. The timings must satisfy leaseDurationSeconds > renewDeadlineSeconds > 1.2 x retryPeriodSeconds, or the chart refuses to render. See broker_leader and broker_leader_election_active on /metrics for which replica is leading. Each replica keeps 4 database connections open and grows its pool up to broker.dbPoolMaxSize (default 32) under load. During a rolling update the old and new pods overlap, so the database must accept:
(dbPoolMaxSize is raised to broker.audit.inflightPermits + 8 if it is lower.) With the bundled database the chart does this for you: it leaves PostgreSQL’s default max_connections of 100 while that is enough and raises it otherwise (150 for three replicas), and refuses to render an explicit database.maxConnections below the requirement. With an external database, check its max_connections against the formula before raising replicaCount; helm install prints the number in its notes.

Use External PostgreSQL

For production-grade deployments — managed services like AWS RDS / Cloud SQL / a CloudNativePG cluster you already run — disable the bundled Postgres and point the broker at the external instance:
Create the secret in the kguardian namespace before installing:
Then install:

AI Assistant

The assistant is optional and off by default. It is one workload — the LLM Bridge — that runs the 23 tools and the policy/seccomp generation in-process against the Broker. Create a Secret holding your provider API key under the key api-key, then turn the assistant on with the provider name and the Secret:
The chart writes the right environment variable for the provider you name — OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEY, or GITHUB_TOKEN. For a multi-provider setup, configure each one under llmBridge.secrets.* instead and leave ai.provider empty. Check it came up with the LLM Bridge health endpoint — hasProvider must be true:

Use an OpenAI-Compatible Gateway

To route the assistant through LiteLLM, vLLM, or an enterprise proxy instead of the vendor API, set ai.baseUrl — and, because a gateway’s model ids are its own, usually ai.model too:
Both apply to whichever provider ai.provider names. Setting them without ai.provider fails the install rather than being silently ignored.
You still need an API key. Provider availability is decided by the API key alone, never by the base URL. Point ai.baseUrl at a gateway without setting ai.secret and /health reports hasProvider: false, chat requests fail with 503 No LLM provider configured, and nothing in the error mentions the gateway you configured.Set ai.secret to a Secret holding your gateway’s virtual key — or any non-empty dummy value if the gateway is unauthenticated. This is the most common first-run mistake.
The request path is appended to your base URL verbatim. kguardian never inserts, removes, or rewrites a /v1 segment; only a trailing slash is normalised.If your gateway serves /v1/chat/completions, include /v1 in ai.baseUrl. If it serves /chat/completions, do not. LiteLLM answers on both, which is why a value that works against LiteLLM can return 404 from vLLM or a proxy that exposes only the /v1 form.
The variable each value maps to depends on the provider: Gemini’s asymmetry is deliberate — the key is GOOGLE_API_KEY while the base URL and model are GEMINI_*, matching each vendor’s own convention. ANTHROPIC_BASE_URL is not new: the Anthropic SDK already honoured it, and the chart now sets it explicitly so a typo is reported against the value rather than surfacing as a connection error. Errors are diagnostic rather than silent. A malformed or scheme-less base URL fails before any request is sent, naming the variable. A 404 or 405 from an OpenAI-compatible, Copilot, or Gemini gateway names the exact URL that was called and the variable to fix. A whitespace-only value counts as unset, consistently with the API keys.

Serve the MCP Endpoint

The LLM Bridge can serve its 23 tools to external MCP clients at POST /mcp on its existing port, so Claude Code or another MCP client can query your cluster’s observed telemetry and generate policies from it. It is off by default:
The chart refuses to render an unauthenticated endpoint unless you set ai.mcp.auth.allowUnauthenticated=true deliberately, because the endpoint serves cluster telemetry to anything that can reach the ClusterIP Service. ai.mcp.rateLimitPerMin moves the per-replica ceiling from its default of 300. The full client-side walkthrough — port-forward, claude mcp add, .mcp.json, and what the tools answer — is in Connect an MCP Client.

Gate the UI Behind SSO

The UI and the Broker API are unauthenticated by default. In a Gateway API environment running an identity-aware proxy (oauth2-proxy in front of an OIDC provider such as Dex, Keycloak or Authentik), the chart can put the frontend route behind SSO:
This renders two Envoy Gateway objects, and nothing at all while enabled is false: a SecurityPolicy that ext-auths the frontend route against oauth2-proxy, and an /oauth2/* HTTPRoute so the OIDC callback and /oauth2/userinfo are served same-origin. The UI detects the session from /oauth2/userinfo at runtime, so there is no rebuild and no build flag; without the proxy it stays in local mode. If SSO is on but the UI still shows local mode, check where /oauth2/userinfo is answered: curl -sI https://<ui-host>/oauth2/userinfo. A 204 with X-Kguardian-Sso: none came from the UI’s own server, so the /oauth2/* route is missing or not attached and the request never reached oauth2-proxy. oauth2-proxy itself answers 200 with a session or 401 without one.
The oauth2-proxy Service usually lives in another namespace, so a ReferenceGrant in that namespace must allow SecurityPolicy and HTTPRoute from the release namespace to reference it. The chart cannot render this for you. Without it the policy attaches but the route cannot reach the proxy.
This gates the UI route only. The Broker API is a separate Service. To authenticate it too, turn on scoped broker tokens (broker.auth.enabled). The frontend’s /api proxy then adds the read token on the server side and drops the SSO Authorization header before the request reaches the broker. SSO also doesn’t cover the MCP endpoint above, which has its own bearer token. SSO guards the Gateway route, not the frontend Service itself. With broker auth on, anything in the cluster that can reach the frontend Service directly can read everything the broker holds, because the proxy adds the read token for it. Restrict ingress to the frontend pods with a NetworkPolicy that admits only your Gateway.

Restrict the UI’s Host Names

The UI server answers only the host names the chart knows: frontend.ingress.hosts (and the TLS hosts) when the chart’s ingress is on, frontend.sso.hostnames when SSO is on, the frontend Service’s in-cluster DNS names (the full one ends in global.clusterDomain, default cluster.local), and anything in frontend.allowedHosts. Any other Host header gets a 403, so a web page cannot rebind its own name to the UI and use the /api proxy, and the read token it adds, from a user’s browser. localhost and IP addresses are always answered, so kubectl port-forward keeps working.
A user who opens the UI on a name the chart does not know sees Blocked request. This host ("…") is not allowed. (the hint that follows, about vite.config.js, does not apply: set frontend.allowedHosts). The chart cannot see these names, so list them yourself:
  • an nginx nginx.ingress.kubernetes.io/server-alias annotation;
  • a LoadBalancer or NodePort Service reached by a DNS name;
  • hostnames on your own HTTPRoute beyond frontend.sso.hostnames;
  • the name a controller rewrites Host to (nginx upstream-vhost, an Istio authority rewrite);
  • the load balancer’s own DNS name, such as an ALB’s internal-k8s-….elb.amazonaws.com.
frontend.allowedHosts also accepts one comma-separated string. A wildcard is broader than in Kubernetes: *.corp.example.com matches corp.example.com itself and subdomains at any depth, not just one label. With none of these set, or with an ingress rule whose host is empty (a catch-all) and no frontend.allowedHosts, the server answers every host, as earlier releases did, and logs a warning at startup.

Install the CLI Plugin

After deploying the controller, install the CLI to generate policies.

Verification

After installation, verify all components are working:

1. Check Pod Status

2. Check Controller Logs

3. Check Broker Connectivity

4. Test CLI Connection

If all checks pass, kguardian is ready to use!

Upgrade

To upgrade kguardian to a newer version:
Always review the Release Notes before upgrading, especially for breaking changes in major versions.
Upgrading the Broker past 1.19.4 with a bloated pod_compute_latest. The first Broker after 1.19.4 truncates pod_compute_latest in a migration that gives up on its table lock after 5 seconds. If the table is already tens of GB, the old Broker’s queries on it can outlast that on every attempt, and the new Broker pod crash-loops while the old one keeps serving. Before upgrading, run the truncate yourself from psql, retrying until it succeeds (the table is a live cache and refills within one sample interval):
Or scale the old Broker to zero for the moment the new one migrates. Either way, the migration then finds an empty table.
Upgrading from an older Broker with large tables. A few migrations add indexes to tables that grow large: pod_traffic and pod_syscalls (2026-06-01, 2026-07-06, 2026-09-26) and image_sbom_components (2026-09-28). When the table is under 256 MiB the migration builds the index itself. When it is larger the migration skips the build, because a plain CREATE INDEX blocks every insert into the table for as long as it runs and could outlast the new pod’s liveness probe. The leader Broker then builds the same index with CREATE INDEX CONCURRENTLY in the background, which blocks no writes; the log shows building index CONCURRENTLY and index built. Until it finishes, reads that use the index are as slow as they were before the upgrade. The compute history indexes are built the same way, on every install.

Uninstall

To completely remove kguardian:
This will not remove generated policies that were already applied to your namespaces. Clean those up separately if needed.

Quick Start

Jump right into generating policies

Connect an MCP Client

Point Claude Code or another MCP client at your cluster

Troubleshooting

Fix common installation issues

Architecture

Understand how components work