Available with limitations in: Open/CE, Core, BE, SE, SE+
Available without limitations in: Ultimate/EE
Included in extensions: Advanced Infrastructure Security
The module lifecycle stage: General Availability
The admission-policy-engine module implements support for admission security policies in a Kubernetes cluster.
Admission policies are rules applied to objects (e.g., Pod and Service) at the time of their creation or modification in the cluster (but not during their operation), based on the information provided in their manifest. These policies are aimed at formalizing parameters that are allowed or prohibited in object manifests.
Policies are divided into three categories:
- Pod Security Standards: Policies that comply with the relevant Pod Security Standards.
- Operational policies: Policies for creating additional requirements for objects by validating the values of parameters that are not directly related to security (for example, a list of allowed prefixes for container images, an image download policy, a list of required container images, etc.).
- Security policies: Policies for creating additional requirements on objects by validating the values of security-related parameters (e.g., container access to the host’s IPC or PID namespaces, privilege lists for containers, etc.).
These policies complement each other. If multiple policies are applied to a single namespace, objects are validated against each of them. If even one policy is violated, the object will not be created.
In addition to policies that prohibit using parameters different from the set requirements, the module supports the SecurityPolicyException resource, which allows creating fine-grained exceptions from security policy checks. With this resource, you can allow using specific parameters for individual pods or containers without changing security policies applied to the entire namespace.
How validation failure messages are displayed
Depending on how pods are created, there are differences in how the API generates messages regarding validation failures (violations of established policies):
- If a pod is created directly, the validation error is returned in the API response indicating a validation failure (policy violation).
- If pods are created via Deployment, the required number of ReplicaSets is created, which in turn attempt to create the pods. In this case, the validation error is not returned in the API response but is displayed in the namespace events or the corresponding ReplicaSet events.
Pod validation when policies are modified or new ones are added
For all three policy categories (Pod Security Standards, operational, and security policies), there is no provision for automatically recreating existing pods when changing existing policies or adding new ones. Pods that existed prior to changes being made to the policy in use or prior to a new policy being added will continue to run until they are restarted. Upon restart, they will be validated against the new rules.
The admission-policy-engine module provides alerts (kind: ClusterObservabilityAlert) for such cases, notifying you of pods in the namespace that violate policies after an existing policy is modified or a new one is added.
To get a list of alerts, use the command:
d8 k get clusterobservabilityalerts
Output example:
NAME SEVERITY STATUS DURATION SUMMARY AGE
SecurityPolicyViolation-f3a77d1dd2175402-1777370195 1 Firing 5h Alerting PrometheusUnavailable 5h1m
OperationPolicyViolation-9b21d0c871796913-1777370435 1 Firing 6h Alerting PrometheusUnavailable 6h1m
To view information about a specific alert, use the following command:
d8 k get clusterobservabilityalert OperationPolicyViolation-9b21d0c871796913-1777370435 -oyaml
Pod Security Standards
Pod Security Standards (PSS) is an official Kubernetes standard that defines three security levels for pods, limiting their privileges. Restrictions are enforced by prohibiting the setting of certain parameters in the pod manifest.
A layered structure is used — each higher level of protection uses all the rules of the previous level and adds its own.
The following protection levels are regulated:
Privileged: An unrestricted policy with the widest possible level of permissions (no restrictions).Baseline: A minimally restrictive policy that prevents the most known and popular ways of privilege escalation. Allows using the standard (minimally specified) pod configuration.Restricted: A policy with significant restrictions. Imposes the strictest requirements on pods.
In the Deckhouse Platform, these policies are implemented using Gatekeeper and enforced by the admission controllers of the admission-policy-engine module, rather than the Kubernetes Pod Security Admission controller. Only the policy descriptions are taken from Kubernetes.
You can read more about each set of policies and their restrictions in the Kubernetes documentation.
Configuring PSS policies for namespaces is done by setting a special label security.deckhouse.io/pod-policy=<POLICY_NAME> on the corresponding namespace.
The default policy can be overridden globally (in the module settings).
The module does not apply policies to system namespaces.
When the multitenancy-manager module is enabled, it creates its own OperationPolicy objects (for example, in the default namespace). These are not affected by the podSecurityStandards settings.
Example of setting the Restricted policy for all pods in the my-namespace namespace:
d8 k label ns my-namespace security.deckhouse.io/pod-policy=restricted
Additionally, it is possible to configure the policy enforcement mode. The following modes are supported:
deny: Prohibit starting pods that do not satisfy the policy.warn: Start pods that do not satisfy the policy, but issue a warning.dryrun: Start pods that do not satisfy the policy, do not issue a warning to the user, but record violations in security reports.
Configuring the policy enforcement mode is done by setting the label security.deckhouse.io/pod-policy-action=<POLICY_ACTION> on the corresponding namespace.
To set the policy enforcement mode globally, use the enforcementaction parameter.
Example of setting the “warn” mode for PSS policies for all pods in the my-namespace namespace:
d8 k label ns my-namespace security.deckhouse.io/pod-policy-action=warn
Operational policies
Operational policies are rules aimed at achieving application security best practices, but not directly related to the validation of classic security-related parameters (for example, a list of allowed prefixes for container images, an image download policy, a list of required container images, etc.).
Operational policies are described using the OperationPolicy custom resource.
In this resource, each parameter is responsible for a separate check applied to resources.
Using the OperationPolicy custom resource allows you to define additional requirements for the resources being created (high-level declarative operational policies) without explicitly interacting with Gatekeeper.
We recommend setting the following minimum set of operational policies:
apiVersion: deckhouse.io/v1alpha1
kind: OperationPolicy
metadata:
name: common
spec:
enforcementAction: Deny
policies:
allowedRepos:
- myrepo.example.com
- registry.deckhouse.ru
requiredResources:
limits:
- memory
requests:
- cpu
- memory
disallowedImageTags:
- latest
requiredProbes:
- livenessProbe
- readinessProbe
maxRevisionHistoryLimit: 3
imagePullPolicy: Always
priorityClassNames:
- production-high
- production-low
checkHostNetworkDNSPolicy: true
checkContainerDuplicates: true
match:
namespaceSelector:
labelSelector:
matchLabels:
custom-operation-policy/enabled: "true"
Policy application is implemented through settings located in the spec.match parameter.
When specifying:
match:
namespaceSelector:
labelSelector:
matchLabels:
custom-operation-policy/enabled: "true"
To apply the above policy, it is sufficient to add the label custom-operation-policy/enabled: "true" to the desired namespace.
Unlike PSS, the label name can be anything. Only a match between the label in the policy selector and the corresponding namespace is required.
You can read more detailed information about using selectors in the selector setup description.
It is also possible to specify the action to be applied for the policy.
The spec.enforcementAction parameter is used for this.
The following modes are supported:
Deny: Prohibit starting pods that do not satisfy the policy.Warn: Start pods that do not satisfy the policy, but issue a warning.Dryrun: Start pods that do not satisfy the policy, do not issue a warning to the user, but record violations in security reports.
Based on this example, you can create your own policy with the necessary settings.
Security policies
Security policies are rules aimed at achieving application security best practices by validating the values of security-related parameters.
Security policies are described using the SecurityPolicy custom resource.
In this resource, each parameter is responsible for a separate check applied to resources.
Using this resource, it is possible to construct a security policy similar to a PSS policy of any level.
Using the custom SecurityPolicy resource allows you to define additional requirements for the resources being created (high-level declarative security policies) without explicitly interacting with Gatekeeper.
Example of a security policy:
apiVersion: deckhouse.io/v1alpha1
kind: SecurityPolicy
metadata:
name: mypolicy
spec:
enforcementAction: Deny
policies:
allowHostIPC: true
allowHostNetwork: true
allowHostPID: false
allowPrivileged: false
allowPrivilegeEscalation: false
allowedFlexVolumes:
- driver: vmware
allowedHostPorts:
- max: 4000
min: 2000
allowedProcMount: Unmasked
allowedAppArmor:
- unconfined
allowedUnsafeSysctls:
- kernel.*
allowedVolumes:
- hostPath
- projected
fsGroup:
ranges:
- max: 200
min: 100
rule: MustRunAs
readOnlyRootFilesystem: true
requiredDropCapabilities:
- ALL
runAsGroup:
ranges:
- max: 500
min: 300
rule: RunAsAny
runAsUser:
ranges:
- max: 200
min: 100
rule: MustRunAs
seccompProfiles:
allowedLocalhostFiles:
- my_profile.json
allowedProfiles:
- Localhost
supplementalGroups:
ranges:
- max: 133
min: 129
rule: MustRunAs
match:
namespaceSelector:
labelSelector:
matchLabels:
security-policy: mypolicy
The allowPrivilegeEscalation and allowPrivileged parameters default to false — even if not explicitly specified. This means that containers will not be able to run in privileged mode or escalate privileges. To allow such behavior, set the parameter to true.
Policy application is implemented through settings located in the spec.match parameter.
When specifying:
match:
namespaceSelector:
labelSelector:
matchLabels:
security-policy: mypolicy
To apply the above policy, it is sufficient to add the label security-policy: mypolicy to the desired namespace.
Unlike PSS, the label name can be anything. Only a match between the label in the policy selector and the corresponding namespace is required.
You can read more detailed information about using selectors in the selector setup description.
It is also possible to specify the action to be applied for the policy.
The spec.enforcementAction parameter is used for this.
The following modes are supported:
Deny: Prohibit starting pods that do not satisfy the policy.Warn: Start pods that do not satisfy the policy, but issue a warning.Dryrun: Start pods that do not satisfy the policy, do not issue a warning to the user, but record violations in security reports.
Security policy exceptions
SecurityPolicyException is a resource that lets you create fine-grained exceptions from security policy checks for individual pods and containers. It allows you to avoid excluding an entire namespace from checks and instead define only the necessary exceptions from a specific rule for a pod or container.
Adding exceptions
To add exceptions for a pod or container, do the following:
-
Create a SecurityPolicyException object describing the required exceptions.
It is recommended that you describe the reason for each exception in the rule’s
metadatafield (for example,metadata.description). This makes auditing and maintenance easier. -
In the pod template (usually via
spec.template.metadata.labelsin a Deployment, StatefulSet, or DaemonSet resource), add one of the following labels referencing the exception:security.deckhouse.io/security-policy-exception: <exception-name>: Exception for the entire pod.security.deckhouse.io/security-policy-exception.container.<container-name>: <exception-name>: Exception for a specific container.
Priority when selecting an exception for a container:
- The label
security.deckhouse.io/security-policy-exception.container.<container-name>is checked first. - If the container-specific label is absent, the exception from
security.deckhouse.io/security-policy-exceptionis used.
If a container-specific label is set for a container but it points to an invalid or non-existent SecurityPolicyException object, it still has priority over the global label and may lead to pod placement denial.
Configuration example
For this example, consider a pod that requires:
- Permission to use the
hostNetworkparameter for the entire pod. - Permission to use the
privilegedparameter only for thesample-initcontainer.
Without the SecurityPolicyException resource, allowing these parameters would require implementing a custom security policy where these settings could be allowed for any pod in the cluster.
With SecurityPolicyException, it is enough to create only the following resources:
-
Exception to allow the
hostNetworkparameter:apiVersion: deckhouse.io/v1alpha1 kind: SecurityPolicyException metadata: name: allow-hostnetwork-pod spec: network: hostNetwork: allowedValue: true metadata: description: >- Pod requires host network mode for node-level network diagnostics. -
Exception to allow the
privilegedparameter in thesample-initcontainer:apiVersion: deckhouse.io/v1alpha1 kind: SecurityPolicyException metadata: name: allow-privileged-init-container spec: securityContext: privileged: allowedValue: true metadata: description: >- Container init requires privileged mode to access host-level networking features.
After that, the corresponding labels need to be added to the pod template:
apiVersion: apps/v1
kind: Deployment
metadata:
name: example
spec:
template:
metadata:
labels:
# General exception applicable to the entire pod.
security.deckhouse.io/security-policy-exception: allow-hostnetwork-pod
# Exception applicable to the sample-init container.
security.deckhouse.io/security-policy-exception.container.sample-init: allow-privileged-init-container
spec:
hostNetwork: true
...
containers:
- name: sample-init
securityContext:
privileged: true
Modifying Kubernetes resources
The module allows you to use the Gatekeeper Custom Resources to modify objects in the cluster, such as:
- AssignMetadata — defines changes to the
metadatasection of a resource. - Assign — any change outside the
metadatasection. - ModifySet — adds or removes entries from a list, such as the arguments to a container.
- AssignImage — to change the
imageparameter of the resource.
You can read more about the available options in the gatekeeper documentation.
Availability of the module components
The module is on the critical path of the cluster: while its validating webhook is unavailable, the API server rejects the requests the webhook intercepts.
Why the validating webhook is a critical component
The gatekeeper-controller-manager deployment serves the ValidatingWebhookConfiguration named d8-admission-policy-engine-config. Its main webhook is configured with failurePolicy: Fail, which means that a request the API server cannot deliver to the webhook is rejected rather than admitted.
While the webhook is unavailable, no object bypasses the policies, but the objects the webhook intercepts cannot be created or changed during that time either.
While no replica of gatekeeper-controller-manager is available, the following stops working in the namespaces the webhook covers:
- Creation, modification and deletion of pods, as well as
d8 k exec,d8 k attachandd8 k debug - Creation, modification and deletion of Role, RoleBinding and Gatekeeper constraints
d8 k execandd8 k attachin namespaces whose names start withd8-andkube-, which a separate webhook withfailurePolicy: Failintercepts
The mutating webhook is configured differently: its failurePolicy is Ignore, and an unavailable deployment only means that mutations are not applied.
Disabling the module unblocks the cluster, but the control goes with it: no policy is enforced any more, and an object a policy used to forbid is created without hindrance. The FAQ describes what to do while the webhook is unavailable.
Which objects are excluded from validation
The exclusions define which objects can be created and changed while the webhook is unavailable.
The webhooks of the module do not validate the following:
- Objects in namespaces with the
heritage: deckhouselabel, which the main webhook excludes throughnamespaceSelector - Objects with the
gatekeeper.sh/operation: webhooklabel that Gatekeeper sets on its own pods, which the webhooks exclude throughobjectSelector - Requests from service accounts of the
d8-virtualizationnamespace
Namespaces that also carry the security.deckhouse.io/enable-security-policy-check label, including d8-admission-policy-engine, are validated by a separate webhook.
The exclusion by the gatekeeper.sh/operation: webhook label is required in the namespace of the module. Without this exclusion, the webhook could not recover from its own outage. Once the last replica is gone, the API server would reject the creation of a replacement pod, since no replica is available to validate it.
d8 k exec and d8 k attach are not covered by the exclusion by label. The object of such a request cannot have labels, and the API server never calls a webhook with objectSelector for it. In the other namespaces these requests are intercepted by separate webhooks without objectSelector, which select the same namespaces as the main webhook and the webhook for namespaces with the security.deckhouse.io/enable-security-policy-check label.
The exclusions do not apply to the webhook that intercepts d8 k exec and d8 k attach in the d8-* and kube-* namespaces: it has neither namespaceSelector nor objectSelector, so d8 k exec and d8 k attach into the pods of the module are blocked during an outage as well.
Third-party components
List of third-party software used in the admission-policy-engine module:
-
Gatekeeper 3.22.2
License: Apache License 2.0
Policy-based control for cloud native environments