Kubernetes gives teams tremendous freedom. That freedom becomes a liability when someone deploys a privileged container, uses an uncontrolled image, omits required ownership metadata, or creates workloads that violate your platform standards.
OPA Gatekeeper lets you express those controls as policy and evaluate Kubernetes resources at admission time, before unwanted configuration is persisted in the cluster.
In this guide, we’ll install Gatekeeper, create a policy using Rego v1,
instantiate it with a Constraint, test it locally with Gator, introduce
policies through dryrun, warn, and deny, organize policies for
GitOps, handle exceptions, monitor Gatekeeper, and cover the operational
decisions required for production.
This tutorial focuses on validation policies. Gatekeeper also supports mutation, but mutation is outside the scope of this guide.
Why Policy-as-Code?
Kubernetes RBAC answers:
Who may perform an operation?
Gatekeeper answers:
Is the requested resource configuration allowed?
They complement each other.
Challenge RBAC Gatekeeper
Block privileged Cannot inspect the Validate
containers requested Pod securityContext
specification
Require resource limits Cannot enforce workload Validate resource fields configuration
Restrict image Cannot validate image Validate image registries names references
Require ownership Cannot require Validate required labels arbitrary metadata labels
Restrict workload Controls API Evaluates the submitted configuration permissions object
Gatekeeper can evaluate resources during admission and can also audit resources that already exist.
Conceptually:
Developer / Controller
|
v
Kubernetes API Server
|
v
Gatekeeper validation
|
+---- compliant ----> resource accepted
|
+---- violation ----> deny / warn / dryrunThis gives platform teams a policy layer that is separate from application code and separate from Kubernetes RBAC.
Architecture Overview
Gatekeeper uses Kubernetes-native custom resources to define and instantiate policy.
The diagram could not be displayed. Its source is available below.
Diagram source
flowchart LR
U[kubectl / Controller]
--> API[Kubernetes API Server]
--> WH[Gatekeeper Admission Webhook]
--> PE[Policy Evaluation]
CT[ConstraintTemplates] --> PE
C[Constraints] --> PE
PE -->|allow / deny / warn| API
AUDIT[Gatekeeper Audit]
--> PE
--> V[Constraint Status / Metrics]The main validation components are:
- ConstraintTemplate — defines reusable policy logic and the parameter schema.
- Constraint — instantiates a template and specifies where and how the policy applies.
- Admission webhook — evaluates matching admission requests.
- Audit controller — periodically evaluates existing resources and reports violations.
- OPA/Rego — provides the policy language and evaluation engine used by Rego-based templates.
A useful mental model is:
ConstraintTemplate = function definition
Constraint = function invocation + scope + enforcement modeStep 1: Install Gatekeeper
You can install Gatekeeper using its published manifests or Helm.
Using kubectl
For a pinned Gatekeeper release:
kubectl apply -f \
https://raw.githubusercontent.com/open-policy-agent/gatekeeper/v3.22.2/deploy/gatekeeper.yamlPin the version in automation rather than referencing an unversioned branch.
Wait for the deployment:
kubectl get pods -n gatekeeper-system -wThen verify the webhook:
kubectl get validatingwebhookconfigurations | grep gatekeeperUsing Helm
For production, Helm makes installation settings easier to manage declaratively.
Add the repository:
helm repo add gatekeeper \
https://open-policy-agent.github.io/gatekeeper/charts
helm repo updateInstall Gatekeeper:
helm install gatekeeper gatekeeper/gatekeeper \
--namespace gatekeeper-system \
--create-namespace \
--set replicas=3 \
--set auditInterval=300 \
--set constraintViolationsLimit=100Verify:
kubectl get pods -n gatekeeper-system
kubectl get validatingwebhookconfigurations | grep gatekeeperFor a real production environment, store the Helm values in Git rather
than maintaining a long --set command.
Step 2: Understand ConstraintTemplates
A ConstraintTemplate defines reusable policy logic.
Modern Gatekeeper can evaluate policies using different engines. This guide uses Rego v1 explicitly.
A template contains:
- The name of the generated Constraint kind.
- The schema for parameters accepted by that Constraint.
- The policy code.
Here is a minimal template requiring labels.
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
name: k8srequiredlabels
annotations:
description: "Requires resources to contain specified labels"
spec:
crd:
spec:
names:
kind: K8sRequiredLabels
validation:
openAPIV3Schema:
type: object
properties:
labels:
type: array
items:
type: string
targets:
- target: admission.k8s.gatekeeper.sh
code:
- engine: Rego
source:
version: "v1"
rego: |
package k8srequiredlabels
violation contains {"msg": msg} if {
required := input.parameters.labels[_]
not object.get(
input.review.object.metadata,
"labels",
{}
)[required]
msg := sprintf(
"Resource is missing required label: %s",
[required]
)
}Notice the explicit:
code:
- engine: Rego
source:
version: "v1"This opts the template into Rego v1 syntax.
The rule:
violation contains {"msg": msg} if {adds a violation when its conditions are true.
input.review.object contains the Kubernetes object being evaluated,
while input.parameters contains values supplied by the Constraint.
Step 3: Create Your First Policy
We’ll require Deployments to have:
owner
teamlabels.
Save the previous template as:
constraint-template.yamlApply it:
kubectl apply -f constraint-template.yamlVerify that Gatekeeper generated the new Constraint CRD:
kubectl get constrainttemplates
kubectl get crd | grep k8srequiredlabelsCreate the Constraint
Now instantiate the template:
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
name: require-workload-ownership
spec:
enforcementAction: dryrun
match:
kinds:
- apiGroups:
- apps
kinds:
- Deployment
excludedNamespaces:
- kube-system
- gatekeeper-system
parameters:
labels:
- owner
- teamApply it:
kubectl apply -f constraint.yamlWe deliberately start with:
enforcementAction: dryrunrather than immediately blocking resources.
Step 4: Test the Policy
Create a non-compliant Deployment:
kubectl create deployment nginx \
--image=nginx:1.27Because the policy is currently in dryrun, the Deployment is still
created.
Inspect the Constraint:
kubectl get k8srequiredlabels \
require-workload-ownership \
-o yamlLook under:
status:
violations:Gatekeeper’s audit controller will report existing objects that violate the Constraint.
Now create a compliant manifest:
apiVersion: apps/v1
kind: Deployment
metadata:
name: nginx-compliant
labels:
owner: platform
team: engineering
spec:
replicas: 1
selector:
matchLabels:
app: nginx-compliant
template:
metadata:
labels:
app: nginx-compliant
spec:
containers:
- name: nginx
image: nginx:1.27Apply it:
kubectl apply -f nginx-compliant.yamlThe Deployment itself contains the required labels and therefore satisfies the policy.
Step 5: Test Policies Locally with Gator
Do not rely exclusively on testing policies against a live cluster.
Gator lets you evaluate Gatekeeper templates, constraints, and Kubernetes manifests locally and is suitable for CI pipelines.
Install it, for example on macOS:
brew install gatorA useful repository layout is:
test/
├── template.yaml
├── constraint.yaml
└── samples/
├── allowed.yaml
└── denied.yamlUse Gator’s verification workflow to test policy behavior before deployment:
gator verify test/The exact test-suite structure depends on how you organize your Gator tests, but the important workflow is:
The diagram could not be displayed. Its source is available below.
Diagram source
flowchart LR
A[Edit Policy]
--> B[Gator Verification]
--> C[Pull Request]
--> D[CI]
--> E[Cluster dryrun]
--> F[warn]
--> G[deny]Policy code should go through the same review and automated validation process as application or infrastructure code.
Step 6: Introduce Enforcement Gradually
A policy that is logically correct can still break existing workloads.
Use Gatekeeper’s enforcement modes as a rollout mechanism.
Dry-Run
Start with:
spec:
enforcementAction: dryrunGatekeeper records violations without blocking admission.
Use this phase to answer:
- How many resources violate the policy?
- Which teams are affected?
- Are system workloads affected?
- Does the policy have false positives?
- Do legitimate exceptions exist?
Inspect violations:
kubectl get constraints -o json | jq '
.items[] |
{
name: .metadata.name,
violations: .status.totalViolations,
auditTimestamp: .status.auditTimestamp
}
'Warn
After remediation, change to:
spec:
enforcementAction: warnDevelopers receive feedback during admission, but the operation continues.
For example:
Warning: [require-workload-ownership]
Resource is missing required label: ownerThis phase makes policy impact visible directly in developer workflows.
Deny
Once the policy is proven and existing violations have been addressed:
spec:
enforcementAction: denyNow non-compliant matching requests are rejected.
The production lifecycle should therefore normally be:
dryrun
|
v
observe + remediate
|
v
warn
|
v
developer feedback
|
v
denyDo not skip directly from policy creation to cluster-wide denial unless the control and its blast radius are already well understood.
Step 7: Build Common Policies
Writing policies yourself is useful for understanding Gatekeeper.
For common Kubernetes controls, however, review the official Gatekeeper policy library before maintaining a custom implementation. Reusable policies should be treated as versioned platform components.
The following examples are intentionally compact and illustrate policy patterns. They should be validated against your organization’s requirements before production use.
Block Privileged Containers
A simple policy can reject containers explicitly configured with:
securityContext:
privileged: trueapiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
name: k8sblockprivilegedcontainers
spec:
crd:
spec:
names:
kind: K8sBlockPrivilegedContainers
targets:
- target: admission.k8s.gatekeeper.sh
code:
- engine: Rego
source:
version: "v1"
rego: |
package k8sblockprivilegedcontainers
violation contains {"msg": msg} if {
container := input.review.object.spec.containers[_]
object.get(
object.get(container, "securityContext", {}),
"privileged",
false
)
msg := sprintf(
"Privileged container not allowed: %s",
[container.name]
)
}For controller objects such as Deployments, you would evaluate:
spec.template.spec.containersrather than spec.containers.
In production, do not treat this one rule as complete Pod hardening. Use policies aligned with your Kubernetes Pod Security Standard requirements and account for init containers, ephemeral containers, host namespaces, capabilities, privilege escalation, user IDs, seccomp, and related controls.
Require Resource Configuration
A policy can require resource settings on containers.
For example, the policy logic can test:
object.get(container.resources, "limits", {})and verify required CPU or memory fields.
Before enforcing a universal CPU-limit rule, define your platform’s actual resource-management policy. Some organizations deliberately require memory limits and CPU requests while avoiding CPU limits for particular latency-sensitive workloads.
Policy should encode an intentional platform standard, not merely the existence of a Kubernetes field.
Restrict Image Registries
If you use string-prefix matching, call the parameter what it really represents.
For example:
validation:
openAPIV3Schema:
type: object
properties:
allowedPrefixes:
type: array
items:
type: stringThen:
allowed_image(image) if {
prefix := input.parameters.allowedPrefixes[_]
startswith(image, prefix)
}A Constraint could define:
parameters:
allowedPrefixes:
- "registry.internal.company.com/"
- "ghcr.io/my-org/"Using a trailing / is important for simple prefix-based matching.
For stronger controls, parse image references or use a maintained policy implementation rather than assuming arbitrary string prefixes represent registry boundaries.
Block the latest Tag
A policy can reject both:
nginx:latestand an untagged image such as:
nginxbecause an omitted tag resolves to latest.
When implementing this yourself, also account for:
registry.example.com:5000/team/image:1.0
image@sha256:...so that a registry port is not mistaken for an image tag and digest-pinned images remain valid.
For high-assurance supply-chain controls, consider requiring immutable
digests rather than merely banning latest.
Step 8: Advanced Policies Using Cluster Inventory
Some policies can evaluate a resource in isolation:
Does this Deployment have an owner label?Others require knowledge about other resources:
Does a matching PodDisruptionBudget already exist?
Is this Ingress hostname already used?
Does this object reference an approved resource?These are referential policies.
Gatekeeper exposes synchronized cluster data to Rego under:
data.inventoryWhy the PDB Example Is More Complex
It is not sufficient to check:
data.inventory.namespace[namespace]["policy/v1"]["PodDisruptionBudget"][_]That only proves that some PDB exists in the namespace.
A correct policy must establish that the PDB actually applies to the workload being evaluated.
Conceptually:
Deployment labels
|
v
Compare against PDB selector
|
+-- matching PDB -> compliant
|
+-- no matching PDB -> violationThat requires:
- synchronizing the required resource type into Gatekeeper inventory,
- reading the correct namespace inventory,
- evaluating the PDB selector,
- comparing it to the target workload labels,
- handling selector semantics correctly.
Because referential policies introduce additional state and complexity, keep them separate from basic object-validation policies and test them explicitly with synchronized-data fixtures.
For production, prefer a tested policy from your maintained policy library over an oversimplified “PDB exists somewhere in the namespace” check.
Step 9: Organize Policies for GitOps
Treat policies as source code.
A practical repository structure is:
policies/
├── templates/
│ ├── security/
│ │ ├── block-privileged.yaml
│ │ └── require-nonroot.yaml
│ │
│ ├── reliability/
│ │ ├── require-resources.yaml
│ │ └── require-probes.yaml
│ │
│ └── operations/
│ ├── require-labels.yaml
│ └── allowed-registries.yaml
│
├── constraints/
│ ├── production/
│ │ └── strict-policies.yaml
│ │
│ ├── staging/
│ │ └── warn-policies.yaml
│ │
│ └── development/
│ └── dryrun-policies.yaml
│
└── tests/
├── security/
├── reliability/
└── operations/The policy delivery lifecycle becomes:
The diagram could not be displayed. Its source is available below.
Diagram source
flowchart LR
A[ConstraintTemplate]
--> B[Constraint]
--> C[Gator Tests]
--> D[Pull Request]
--> E[CI]
--> F[GitOps]
--> G[dryrun]
--> H[warn]
--> I[deny]Apply templates before constraints:
kubectl apply -f policies/templates/ -R
kubectl apply -f policies/constraints/production/A GitOps controller can manage the same ordering through dependency configuration.
Step 10: Handle Exceptions Safely
Real environments need exceptions.
The important question is not whether exceptions exist, but who is allowed to create them and how they are audited.
Namespace Exclusions
A Constraint can exclude known namespaces:
spec:
match:
excludedNamespaces:
- kube-system
- gatekeeper-systemThis is useful for explicitly managed infrastructure namespaces.
Do not build an ever-growing list of exclusions simply because a workload fails policy. Determine whether the policy or workload should change.
Label-Based Exclusions
Gatekeeper matching can exclude resources based on labels.
For example:
spec:
match:
labelSelector:
matchExpressions:
- key: policy.example.com/exempt
operator: DoesNotExistThis means a resource with:
metadata:
labels:
policy.example.com/exempt: "true"does not match the Constraint.
That is convenient but potentially dangerous.
If workload owners are allowed to add the exemption label themselves, they can bypass the policy themselves.
Use label-based exemptions only when the ability to set the exemption marker is controlled through another trusted mechanism.
Prefer Auditable Exceptions
A production exception process should capture at least:
policy
resource / namespace
owner
business justification
approver
creation date
expiry dateThe desired model is:
Violation
|
+--> Fix workload
|
+--> Fix incorrect policy
|
+--> Approved temporary exception
|
+--> owner
+--> justification
+--> expiryExceptions should be exceptional, reviewable, and removable.
Step 11: Understand Fail-Open vs Fail-Closed
Gatekeeper runs in the Kubernetes admission path.
That makes webhook availability a security and platform-availability decision.
When the admission webhook cannot respond, Kubernetes follows the
webhook’s failurePolicy.
Conceptually:
The diagram could not be displayed. Its source is available below.
Diagram source
flowchart TD
A[Admission Request]
--> B{Gatekeeper available?}
B -->|Yes| C[Evaluate Policy]
B -->|No| D{failurePolicy}
D -->|Ignore| E[Allow request to continue]
D -->|Fail| F[Reject matching request]Fail Open
With:
failurePolicy: Ignorea Gatekeeper outage does not automatically block matching Kubernetes API operations.
The tradeoff is that admission policy may not be enforced while the webhook is unavailable.
Fail Closed
With:
failurePolicy: Failmatching admission requests fail when Gatekeeper cannot evaluate them.
This preserves the enforcement boundary but can affect cluster operations during a Gatekeeper outage.
Choose Deliberately
This is a platform risk decision:
Mode Availability Enforcement during Gatekeeper outage
Fail open Higher May be bypassed Fail closed Depends on webhook availability Preserved
Do not change this behavior casually.
If you require fail-closed enforcement, design Gatekeeper itself for high availability and test failure scenarios such as:
- controller pod loss
- node loss
- network-policy mistakes
- certificate problems
- webhook timeouts
- Gatekeeper upgrades
A security control in the API admission path must have an explicit availability model.
Step 12: Production Installation Considerations
For production, manage Gatekeeper as a platform component.
Version-control:
- Gatekeeper version
- Helm values
- webhook configuration
- replica counts
- resource requests and limits
- audit configuration
- failure-policy decision
- ConstraintTemplates
- Constraints
- exception configuration
- monitoring rules
- upgrade procedure
Scope Policies Carefully
Before introducing a cluster-wide deny, consider the workloads that
may be affected:
- Kubernetes system components
- operators
- GitOps controllers
- CI/CD automation
- monitoring
- storage operators
- service meshes
- certificate controllers
- emergency operational tooling
A policy that blocks ordinary application mistakes can also block a critical operator upgrade if its match scope is too broad.
Keep the Policy Layer Small
Prefer a smaller set of well-understood controls over hundreds of overlapping constraints.
Each enforced policy adds:
- admission behavior
- operational ownership
- developer-facing errors
- testing requirements
- upgrade requirements
- exception requirements
Policy sprawl is still configuration sprawl.
Step 13: Monitoring and Alerting
Gatekeeper exposes Prometheus metrics for both admission and audit behavior.
At minimum, monitor:
- Gatekeeper pod availability
- webhook request errors
- webhook latency
- audit execution
- number of audited violations
The audit metric:
gatekeeper_violationsreports violations found during Gatekeeper’s most recent audit.
Also monitor whether audit itself is running, for example with the Gatekeeper audit-run timestamps exposed by your installed version.
Do not only alert on:
violations > thresholdA broken audit controller can otherwise make the system appear healthy simply because it stopped producing fresh results.
Example ServiceMonitor
If you use Prometheus Operator, create a ServiceMonitor that matches
the labels and metrics port exposed by your Gatekeeper Helm release.
For example:
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: gatekeeper
namespace: gatekeeper-system
spec:
selector:
matchLabels:
gatekeeper.sh/system: "yes"
endpoints:
- port: metricsVerify the actual Service labels and port names first:
kubectl get svc -n gatekeeper-system --show-labels
kubectl get svc -n gatekeeper-system -o yamlDo not assume monitoring selectors remain identical across Gatekeeper releases or installation methods.
Step 14: Troubleshooting
Constraint Is Not Enforcing
Check that the template exists:
kubectl get constrainttemplatesCheck Constraints:
kubectl get constraintsInspect the specific Constraint:
kubectl describe k8srequiredlabels \
require-workload-ownershipVerify:
spec:
enforcementAction:A policy in dryrun intentionally does not block admission.
Also check the Constraint’s match section. The object may simply be
outside the configured scope.
ConstraintTemplate Errors
Inspect the template:
kubectl describe constrainttemplate \
k8srequiredlabelsCheck Gatekeeper logs:
kubectl logs \
-n gatekeeper-system \
-l control-plane=controller-managerFor Rego v1 templates, verify that the policy is under:
code:
- engine: Rego
source:
version: "v1"and uses Rego v1 rule syntax.
Webhook Timeout Errors
Check Gatekeeper pods:
kubectl get pods -n gatekeeper-systemCheck logs:
kubectl logs \
-n gatekeeper-system \
-l control-plane=controller-managerInspect webhook configuration:
kubectl get validatingwebhookconfigurations \
gatekeeper-validating-webhook-configuration \
-o yamlCheck:
failurePolicytimeoutSeconds- namespace selectors
- object selectors
- CA bundle
- webhook Service reference
Then verify the Service and endpoints:
kubectl get svc,endpoints \
-n gatekeeper-systemPolicy Works in Admission but Not Audit
Check the Constraint status:
kubectl get constraints -o yamlVerify the audit controller is running and review its logs.
For referential policies, also verify that required resource types are synchronized into Gatekeeper’s inventory.
Test Before Debugging in the Cluster
Use Gator to reproduce policy evaluation locally:
gator verify test/This helps separate:
policy logic problemfrom:
Gatekeeper / Kubernetes integration problemStep 15: Final Verification
At the end of the rollout, verify the platform in layers.
Gatekeeper
kubectl get pods -n gatekeeper-system
kubectl get validatingwebhookconfigurations | grep gatekeeperTemplates
kubectl get constrainttemplatesConstraints
kubectl get constraintsAudit
Inspect violation counts:
kubectl get constraints -o json | jq '
.items[] |
{
name: .metadata.name,
enforcement: .spec.enforcementAction,
violations: .status.totalViolations
}
'Admission
Test a known non-compliant manifest in a controlled namespace.
For a deny Constraint, it should be rejected.
Then test a compliant manifest and verify that it succeeds.
CI
Run your local policy verification:
gator verify test/Your final policy lifecycle should look like this:
The diagram could not be displayed. Its source is available below.
Diagram source
flowchart LR
A[Define]
--> B[Test]
--> C[Review]
--> D[Dry Run]
--> E[Remediate]
--> F[Warn]
--> G[Enforce]
--> H[Audit]
--> I[Improve]Summary
OPA Gatekeeper adds a policy layer to Kubernetes admission control.
The basic technical model is simple:
ConstraintTemplate
+
Constraint
|
v
Gatekeeper
|
v
Admission + AuditThe harder part is operating policy safely.
A production policy program should therefore do more than write Rego. It should:
- test policies before deployment
- introduce controls through
dryrun,warn, anddeny - scope constraints carefully
- use maintained policy libraries where appropriate
- control and audit exceptions
- understand fail-open versus fail-closed behavior
- monitor both admission and audit health
- version policies and Gatekeeper configuration through GitOps
- treat policy changes as platform changes with a measurable blast radius
The goal is not to reject as many Kubernetes resources as possible.
The goal is to make the safe configuration the predictable configuration, give developers useful feedback when they violate a standard, and prevent known-bad state from reaching production.


