The module lifecycle stage: Experimental
The module has requirements for installation
v1.2.0
Release date: 2026-09-28
The release reduces the log collection load on nodes, drops security events that fail the event registry check, and restricts the module schemas and permissions. The module requires Deckhouse 1.76 and log-shipper 1.0.6 or later.
Highlights
Changes in this release:
- Each log source is read once per node regardless of the number of event codes.
- Events that fail the event registry check are dropped, and the enrichment plugins fill their target fields.
- The module custom resources and configuration are bounded by the schema, the component permissions are narrowed, and destination fields that can contain secrets are hidden from users without access to them.
- A cluster-wide shipper collects pod logs from the namespaces it specifies, and a namespace selector can include and exclude namespaces at the same time.
- A destination that sends events in the CEF format can specify the format version in the header prefix.
New features
This release adds:
- The
D8SecurityEventsManagerAgentDeliveryStalledalert fires when a log-shipper agent has undelivered security events and has not delivered any of them to the gateway for 10 minutes. Previously, such a delay triggered no alert, and the gateway pods remainedReady. The alert specifies the node and the sink. - The
input.kubernetesPods.namespaceSelectorparameter acceptsmatchNamesandexcludeNamestogether: the exclusion is applied after the inclusion. Previously, only one of the lists could be specified. - The
encoding.cef.versionparameter ofClusterSecurityEventDestinationsets the CEF format version:V0writes theCEF:0prefix and is used by default,V1writesCEF:1. The header fields and extensions do not change. The parameter is available for the Kafka, Vector, File, Console, and Socket destinations. - The
spec.cef.versionparameter ofClusterSecurityEventConfigsets the CEF format version for the whole cluster. Theencoding.cef.versionparameter of a destination overrides it. - In DKP 1.78 and later, with the
multitenancy-managermodule enabled,SecurityEventDefinitionandClusterSecurityEventEnrichmentPluginare registered as grantable cluster resources. By default, both resources are available to all projects. To restrict access to them, useClusterResourceGrantPolicy.
Improvements
This release improves:
- Collection configs and destinations are created per shipper item instead of per event code. Previously, a log source with 50 event codes was read 50 times on every node and required a separate disk buffer for each code, up to about 20 GiB per node. Events are still filtered on the node.
- All destinations use one TLS Secret in the
d8-log-shippernamespace instead of a copy per event code. - The gateway determines the event code by the same
extractrules as the agent. A log line that matches several event codes produces an event for each of them. - If a shipper item contains a rule that cannot be converted into a node filter, the logs of the item are sent to the gateway unfiltered and are processed there. This prevents events from being dropped on the node.
- Collection configs and destinations created by earlier versions of the module are removed after the new ones are ready.
Fixes
This release fixes:
- Connections of log-shipper agents to the gateway no longer fail with the
SSL alert number 80error. The error occurred on about every second reconnection and delayed the delivery of events. - The
k8s-pod-info,k8s-container-info, andk8s-nodeuser-infoenrichment plugins fill the target field again. Starting from version 1.0.2, the plugins left the field empty, and the gateway log contained thecjson unavailableerror:actor.id,object.name, andobject.namespaceof events from external modules and static user names were not set. If the JSON library fails to load, the gateway log now contains the reason. - Setting one
encoding.cefparameter inClusterSecurityEventDestinationno longer sets the other parameters to built-in values. Previously,deviceVendor,deviceProduct, anddeviceVersionreceived default values and overrodespec.cefofClusterSecurityEventConfig. - If several
SecurityEventDefinitionobjects declare the same event code and source, the gateway uses the definition with the alphabetically first name. Previously, the severity, category, and description of such events could change between reconciliations. The other definitions are reported by thegateway_config_problematic_cluster_objectsmetric with theduplicate_code_sourceproblem. - The gateway configuration no longer changes between reconciliations when the cluster has several
ClusterSecurityEventConfigobjects. Previously, the order of route conditions and the source of the CEF header defaults depended on the order of objects in the API cache, which caused unnecessary gateway reloads and could change the vendor, product, and version in the CEF header. - The order of parsing rules in the gateway no longer changes between reconciliations. Previously, when two rules covered the same pods or two shippers declared the same source and event code, the applied
parserortransformcould change. - The
cluster-lokidestination is created only when thekube-rbac-proxyCA is available. Previously, the destination could be created without a CA, and all events sent to Loki were lost. parserrules are applied to a shipper that selects pods byinput.kubernetesPods.labelSelector.matchExpressions. Previously, the logs of such a shipper were collected but not parsed, and the fields set intransformremained empty.- A shipper with a pod selector that contains only
NotInandDoesNotExistrequirements covers pods without labels, as Kubernetes and log-shipper do. Previously, the logs of such pods were collected but not parsed. - A
ClusterSecurityEventShipperwithinput.type: KubernetesPodsand namespaces specified ininput.kubernetesPods.namespaceorinput.kubernetesPods.namespaceSelector.matchNamescollects events from these namespaces. Previously, such a shipper remained in theReconcileFailedstate and collected no events. - The controller no longer logs the
failed to remove finalizer during SED cleanupandfailed to update success statuserrors when two reconciliations process the sameSecurityEventDefinitionat the same time. These errors did not affect the result of the reconciliation. - The
gateway.logShipperBuffer.maxEventsparameter is applied. Previously, it was ignored, and a buffer of theMemorytype used the log-shipper default. - TLS Secrets created in
d8-log-shipperby earlier versions of the module are removed when no destination references them. When aSecurityEventDefinitionis deleted, its destination is removed before the Secrets it references. Previously, a reference to a deleted Secret could break the log-shipper configuration for all modules in the cluster. - Changing or removing a shipper item or deleting a shipper removes the related
ClusterLogDestination. Previously, such objects remained in the cluster, were processed by every log-shipper agent, and prevented the shared TLS Secret from being removed. - A shipper with a source that has no
SecurityEventDefinitionno longer blocks the removal of its other objects. Previously, items removed from such a shipper kept their collection configs and continued to collect logs. - The gateway drops events that fail the event registry check. Previously, with the default
Lowseverity threshold, events without a matchingSecurityEventDefinitionor without a required field reached the storage without a severity, category, and description, and severity-based filters and dashboards did not show them. TheD8SecurityEventsManagerGatewayEventRegistryDropsalert fires on such drops.
Security updates
Security updates in this release:
- Destination fields that can contain secrets are marked with
x-kubernetes-sensitive-data: true: the authentication token and password of thehttp,loki, andelasticsearchdestinations, the Kafka SASL password, the Splunk HEC token, and thespec.http.headersvalues. In Deckhouse 1.76 and later, these fields are removed fromget,list, andwatchresponses for users without access to theclustersecurityeventdestinations/sensitivesubresource and are masked in the audit log. In Deckhouse 1.77 and later, they are also masked increate,update, andpatchresponses. The values are stored in etcd unencrypted unlesscontrolPlaneManager.encryptionEnabledis enabled, so use thetokenSecretRefandpasswordSecretReffields instead. - The bundled
cluster-lokidestination verifies the certificate and the host name of Loki. Previously, in Deckhouse versions earlier than 1.76, the verification was disabled, which allowed Loki to be impersonated from the cluster network. - All user-settable fields of the module custom resources and configuration have explicit limits:
maxLengthfor strings,maxItemsfor arrays,maxPropertiesfor maps, andminimumandmaximumfor numbers. SecurityEventDefinition.spec.codeandproduces[].eventCodeof both shipper types must match the^[A-Z0-9_]+$format. Previously, objects with other codes were accepted and ignored. Such objects are now rejected on creation and on aspecchange.- The controller ClusterRole no longer grants cluster-wide access to
secretsandevents. Access to Secrets is limited to the module namespace andd8-log-shipper, and write access in the module namespace is limited togateway-vector-configandgateway-destination-cas. The controller no longer changes Secrets owned by users. - The gateway reloader and the
enrichment-cachesidecar serve metrics only on loopback addresses:127.0.0.1:9256and127.0.0.1:9261. Ports9255and9260, which are reachable from the cluster network, serve only health probes. Prometheus collects the metrics through kube-rbac-proxy. - The
controllerService no longer publishes port8081of the health probes. - RBAC objects are renamed to include the component name:
d8:security-events-manager:controller:rbac-proxy,d8:security-events-manager:aggregator:rbac-proxy,d8:security-events-manager:gateway:rbac-proxy,d8:security-events-manager:gateway:pod-reader, andd8:security-events-manager:controller:log-shipper-secretsfor the Role and RoleBinding ind8-log-shipper.
Breaking changes
Changes that affect backward compatibility:
- The module requires log-shipper 1.0.6 or later.
- The minimum supported Deckhouse version is 1.76. The module is not installed on earlier versions.
- A namespace name in
input.kubernetesPods.namespaceSelectorandinput.kubernetesPods.namespacemust be a valid DNS-1123 label. Previously, invalid names were accepted and broke log collection for the entire shipper. input.kubernetesPods.labelSelector.matchLabelscannot be empty. Previously, an emptymatchLabelswas accepted, and logs were collected from all namespaces without producing events.- A requirement in
input.kubernetesPods.labelSelector.matchExpressionsmust have a non-emptykey, must havevaluesfor theInandNotInoperators, and must not havevaluesfor theExistsandDoesNotExistoperators. Previously, such requirements were accepted, and the shipper remained in theReconcileFailedstate. - The module configuration is bounded by the schema:
replicashas minimum and maximum values, and CPU and memory values must use the plain or suffixed Kubernetes quantity format. The exponential format, for example128e6, is rejected. A ModuleConfig that does not match the schema blocks the module update. SecurityEventDefinition.spec.metadataaccepts onlydocs.desc.{en,ru}andtests.{executionTarget,testability,trigger.commands,trigger.cleanupCommands}. Other keys are removed on the next write of the object.
Upgrade notes
Before upgrading, note the following:
- Update Deckhouse to 1.76 or later before updating the module.
- Update log-shipper to 1.0.6 or later before updating the module. With an earlier log-shipper version, the module keeps the previous collection scheme.
- Before updating, check the ModuleConfig, your
SecurityEventDefinitionobjects with other keys inspec.metadata, and custom resources with long regular expressions, Grok patterns, field paths, or HTTP header values against the new schemas. The bundled destination and event definitions match the new schemas. - Before updating, check shippers for namespace names and
matchExpressionsrequirements that the new schema rejects. Existing objects with such values are not deleted, but any change to them is rejected until the values are corrected. - Find
SecurityEventDefinitionobjects and shippers with event codes that do not match^[A-Z0-9_]+$. They are ignored, and any change to theirspecis rejected. - To read inline destination credentials, grant
get,list, orwatchaccess to theclustersecurityeventdestinations/sensitivesubresource. The module controller and thed8:security-events-manager:admin-kubeconfigrole already have this access. - Add the
security-events-manager.deckhouse.io/credential-secretlabel to manually created Secrets with credentials of alert rule destinations. The controller no longer adds this label. The Secret bundled with the module already has it. - To make destinations created before the update follow
spec.cefofClusterSecurityEventConfig, removedeviceVendor,deviceProduct, anddeviceVersionfrom theirencoding.cef. The values written by earlier versions remain in the objects. - After the update, expect a one-time duplication of events from the collected log files: each file is read again from the beginning of its current rotation.
- After the update, events that fail the event registry check do not reach the storage. If the
D8SecurityEventsManagerGatewayEventRegistryDropsalert fires, find the dropped event code by theenforce_event_registry: droppedline in the log of thevectorcontainer of the gateway, and make sure that the source fills the fields that theSecurityEventDefinitionlists as required.
Docs
Documentation changes:
- The examples page contains a
ClusterSecurityEventShipperthat collects pod logs from specified namespaces, includingmatchNamesandexcludeNamesused together. - The
SecurityEventLoggingTransformationRulesexample is placed in thed8-runtime-audit-enginenamespace, where the shipper that references it runs. The page also states thatspec.selectorof these rules is not used to select pods. - The
k8s-pod-infoexamples use@root.podand@root.namespacefor the pod name and namespace. The previous examples used fields that events do not contain. - The Falco examples select pods by the
app: runtime-audit-enginelabel in thed8-runtime-audit-enginenamespace. The previous examples used a label that module pods do not have. - The path syntax of
transform[].valueandenrich[].args[].valueis described, including brackets and quotes for keys that contain a dot. The Falco examples are corrected accordingly.
v1.1.5
Release date: 2026-09-07
Security updates for the alert-evaluator, controller, and go-hooks images: dependency upgrades remediate known vulnerabilities, and OpenVEX attestations mark unused golang.org/x/crypto packages as not affected.
Highlights
Changes in this release:
- The alert-evaluator, controller, and go-hooks images include updated golang.org/x libraries that remediate High and Critical vulnerabilities found in recent image scans.
- OpenVEX attestations for the controller and module bundle mark vulnerabilities in golang.org/x/crypto/ssh and golang.org/x/crypto/openpgp as not affected, because those packages are not used.
Security updates
Security updates in this release:
- In the alert-evaluator image, golang.org/x/net is updated to v0.56.0, golang.org/x/text to v0.39.0, and golang.org/x/sys to v0.46.0. The update remediates related CVEs, including CVE-2025-47911, CVE-2025-58190, CVE-2026-25680, CVE-2026-25681, CVE-2026-27136, CVE-2026-33814, CVE-2026-39821, CVE-2026-39824, CVE-2026-42502, CVE-2026-42506, CVE-2026-46600, and CVE-2026-56852.
- In the controller and go-hooks images, golang.org/x/crypto is updated to v0.55.0. The update remediates Critical CVE-2026-56854.
- OpenVEX statements for the controller and module bundle mark CVE-2026-56855, CVE-2026-78662, and GO-2026-5932 as not_affected with the vulnerable_code_not_in_execute_path justification. The controller uses golang.org/x/crypto for bcrypt, scrypt, and pbkdf2 through sprig; go-hooks uses it for ocsp and pkcs12 through cfssl. Neither component imports golang.org/x/crypto/ssh or golang.org/x/crypto/openpgp.
v1.1.4
Release date: 2026-09-03
Alert rules are now LogQL, evaluated directly against Loki: ClusterSecurityEventAlertRule is reshaped after PrometheusRule (groups of rules with alert/expr/for/labels/annotations), and the alerting pipeline no longer runs through Vector.
Highlights
Changes in this release:
ClusterSecurityEventAlertRuleis now shaped likePrometheusRule:spec.groups[].rules[], each rule withalert,for,severityLevel,labelsandannotations, and either a hand-written LogQLexpror a structuredmatch/aggregationblock (event codes, sources, categories, actors, grouping, threshold) that compiles into the equivalent LogQL. This is a breaking change to the resource spec — rewrite existing rules before upgrading.- Every rule declares its form in the required
modefield (ExprorMatch), matching the block it sets. The platform web interface usesmodeto show only the fields for the selected form. - A new
alert-evaluatorcomponent, running inside the existingaggregatorDeployment, evaluates each rule’s LogQL query against the cluster’s Loki on a schedule and reports firing rules through the samesecurity_events_alert_last_firing_timestamp_secondsmetric as before. TheD8SecurityEventAlertFiringalert is unchanged. gatewayno longer forwards a copy of matching events toaggregatorover the network:alert-evaluatorreads directly from Loki.aggregator’s TLS material and Vector process are no longer needed.- Alert rules require the
cluster-lokiClusterSecurityEventDestinationto be enabled — without it, LogQL has no data to query.
Improvements
This release improves:
- A rule can now look arbitrarily far back and group by any label LogQL supports, not only the eight fixed fields the previous
aggregation.groupByenum offered. Use a hand-writtenexprfor that. - Each
ClusterSecurityEventAlertRulereports its ownstatus.conditions(Ready/Degraded, with a reason such asLogQLQueryError,LokiDestinationMissingorCardinalityLimitExceeded).
Breaking changes
Changes that affect backward compatibility:
- Every rule in
ClusterSecurityEventAlertRule.spec.groups[].rules[]must declaremode: Exprormode: Match, matching the block it sets (exprormatch). A rule withoutmode, or with amodethat disagrees with its block, is rejected. - The old
spec.match/spec.aggregation/spec.alertfields (one rule per object) are removed. Rewrite each rule underspec.groups[].rules[], using either a hand-written LogQLexpror the samematch/aggregationfields as before (eventCodes,sources,categories,severityMin,actorType,actors,excludeActors,excludeSystemActors,groupBy,threshold,window), now nested per rule. See the module documentation for examples. - The API server now enforces field constraints on
ClusterSecurityEventShipper,PodSecurityEventShipper,ClusterSecurityEventLogTransformation,SecurityEventLogTransformation,SecurityEventDefinition,ClusterSecurityEventEnrichmentPluginandClusterSecurityEventDestinationthat were previously silent. An empty label selector (labelSelector: {}/selector: {}) is now rejected instead of matching everything,ClusterSecurityEventDestination’sspec.*.buffer.maxEventsmust be at least 100, and several free-text fields no longer accept an empty string. - Alert rule evaluation requires the
cluster-lokiClusterSecurityEventDestination(enable it viasecurityEventConfig.destinations). Without it, everyClusterSecurityEventAlertRuleis reportedDegradedand no rule is evaluated; event collection and delivery to other destinations are unaffected.
Upgrade notes
Before upgrading, note the following:
- When rewriting a rule, add
mode: Exprto a rule with a hand-writtenexpr, ormode: Matchto a rule built frommatch/aggregation. - Before upgrading, save every existing
ClusterSecurityEventAlertRule(d8 k get clustersecurityeventalertrule -o yaml). After upgrading, rewrite each rule under the newspec.groups[].rules[]shape — the old top-level spec fields are rejected by the updated CRD schema. - Before upgrading, check whether existing resources violate the newly enforced constraints: run
d8 k get clustersecurityeventshipper,podsecurityeventshipper,clustersecurityeventlogtransformation,securityeventlogtransformation,clustersecurityeventdestination -o yamland look for an emptylabelSelector/selectororbuffer.maxEventsbelow 100. An existing object that violates a constraint keeps working, but the next write to it is rejected. - Confirm
cluster-lokiis enabled (securityEventsManager.securityEventConfig.destinationsincludescluster-loki) before relying on alert rules; otherwise every rule staysDegraded.
Known issues
Known limitations of this release:
- Evaluation state (pending/firing) lives in the evaluator process’s memory: restarting it resets the
forcountdown for every rule. - Grouping by a LogQL label with unbounded cardinality is capped at 500 distinct label sets per rule; the rest are dropped and the rule is reported
CardinalityLimitExceeded, but it keeps firing for the groups it already tracks.
Docs
Documentation changes:
- The alerting section of the examples page is rewritten for the
PrometheusRule-shaped resource, covering both theexprandmatch/aggregationforms, including the line filter that reproduces the oldexcludeSystemActorsbehavior.
v1.1.3
Release date: 2026-08-27
Security fixes for module images: OpenVEX attestations mark golang.org/x/crypto GO-2026-5932 as not affected for the controller and module bundle hooks.
Highlights
Changes in this release:
- Module images publish OpenVEX attestations so CVE scans suppress GO-2026-5932 where openpgp is not on the execute path.
Security updates
Security updates in this release:
- Added OpenVEX statements for GO-2026-5932 (golang.org/x/crypto/openpgp) on the controller and module bundle (go-hooks), status not_affected / vulnerable_code_not_in_execute_path.
- Wired cosign OpenVEX attestation into the module build so CI CVE scans read the statements from the OCI registry.
v1.1.2
Release date: 2026-08-27
Alerting on security events: the new ClusterSecurityEventAlertRule resource counts matching events per key over a time window and raises an ordinary Prometheus alert once the threshold is reached. Adds an Http destination type for webhooks and HTTP collectors, and a new event for containerd integrity violations.
Highlights
Changes in this release:
- New
ClusterSecurityEventAlertRuleresource: describe which events to watch, how to count them, and what the alert should say. - A firing rule raises the Prometheus alert
D8SecurityEventAlertFiring, so it becomes a ClusterAlert object, shows up in the console next to every other alert, keeps its history and travels through the cluster’s existing alert delivery. There is nothing to configure for delivery. - New
Httpdestination type sends security events to any HTTP endpoint — a webhook or an HTTP collector — with authentication, TLS, custom headers and a choice of body framing. - New security event
D8_CONTAINERD_INTEGRITY_VIOLATIONreports a containerd integrity check failure — most often an image refused at pull because its signature did not verify.
New features
This release adds:
ClusterSecurityEventAlertRule(cluster-scoped, short namecsear).spec.matchselects events by event code, minimum severity, category, source component, actor and actor type;spec.aggregationgroups them bygroupByand fires as soon as a group reachesthresholdwithinwindow;spec.alertcarries the summary, description, extra labels and the DKPseverityLevel. A rule withoutaggregationalerts on every matched event.match.excludeSystemActors(enabled by default) keeps platform identities out of a rule, using the same list as the “Exclude system” filter of the console. Routine cluster work dominates the event stream — on an idle cluster the DKP ServiceAccount alone produced 2617 of 2812 secret reads in an hour — so without it the first rule anyone writes is an alert about the platform.match.excludeActorsandmatch.actorstake globs againstactor.idfor anything the built-in list does not cover.- The alert resolves on its own about five minutes after the rule last fired. There is no field to set: the aggregator publishes
security_events_alert_last_firing_timestamp_seconds, the moment a rule last fired, andD8SecurityEventAlertFiringtests the age of it. The metric is exported on the aggregator’s/alert-metrics/metricsendpoint and can be used in dashboards. - New destination type
Http:spec.http.endpointwith optionalauth,tlsandheaders, andframing—JSONArray(default) for receivers that expect one document per request,NewlineDelimitedfor collectors that read NDJSON. The number of delivery attempts for one batch is bounded, so a receiver that keeps rejecting a request it will never accept cannot hold up everything queued behind it: once the attempts run out the batch is dropped into the sink’s error metrics, which the module already alerts on. - New security event
D8_CONTAINERD_INTEGRITY_VIOLATION(categoryRuntime, severityHigh, sourcecontainerd-integrity). It fires when containerd refuses an image whose signature does not verify, and when a file checksum on the node does not match the expected one.object.namecarries the image reference,object.namespacethe namespace it was being pulled into, andmetadata.extrathe message, the file path and the expected and actual hashes where the event provides them. Collected only on nodes running cri ContainerdV2 in the CSE edition; probe traffic from d8-upmeter is filtered out. - kube-audit events now carry
actor.sourceIP, the originating client address taken from the head of the request chain, next to the fullactor.sourceIPschain. Alert rules can group by it (ActorSourceIP), which is how a rule about unauthorized requests identifies the client when there is no identity to group on. - The module runs a new single-replica
alert-aggregatorcomponent that holds the counters, with its own alertD8SecurityEventsManagerAggregatorNotAvailable. While it is down, event collection and delivery continue — events wait in the gateway buffer — but counting windows restart, so a partially reached threshold begins from zero.
Improvements
This release improves:
- Transform source paths can address array elements, for example
sourceIPs[0]. Previously every path segment was treated as a map key, so any array field in a log was unmappable. - Module documentation, CRDs and OpenAPI schemas are now included in the bundle and release images.
Upgrade notes
Before upgrading, note the following:
- No configuration change is required. The alerting components are created on upgrade and stay idle until the first
ClusterSecurityEventAlertRuleexists. Thealert-aggregatorruns a single replica with aRecreaterollout by design — counters live in one process — so its updates and node drains produce a short gap in counting.
Known issues
Known limitations of this release:
- A rule can only match what the cluster’s audit policy records. The default DKP policy logs
listof secrets but not agetof one secret by name, and drops most requests fromsystem:authenticated, so a rule onK8S_SECRET_ACCESSEDwill not seed8 k get secret <name>. Check the audit policy before concluding that a rule does not work. - Counters live in the aggregator’s memory. Restarting it starts every window from zero.
- Grouping by a value with unbounded cardinality is capped: past 500 distinct values for a label the tag is dropped and the alert still fires, one dimension poorer.
- On clusters with
update.blockOnAlertsenabled in the deckhouse module, an alert at or below the configured level (4 by default) blocks DKP releases from being applied while it is firing. TheseverityLeveldefault of 6 stays out of the way; lower it deliberately.
Docs
Documentation changes:
- The examples page gained an alerting section with a rule per detection story: repeated exec into pods by one actor, a burst of RBAC edits grouped per actor and namespace, unauthorized requests grouped by client address, and a containerd integrity violation as the case where a single occurrence is enough. It also gained an HTTP destination example.
v1.0.5
Release date: 2026-08-24
Security fixes for module images: updated Go dependencies in batch hooks and related components to remediate known CVEs.
Highlights
Changes in this release:
- Module images include dependency updates that address known CVEs in Go components.
Security updates
Security updates in this release:
- Bumped golang.org/x/net and golang.org/x/text in hooks/batch and related Go images for CVE remediation.
- CVE-oriented security fixes included in the component images.
v1.0.4
Release date: 2026-08-18
Fixes an upgrade deadlock: the controller could not publish gateway configs and the gateway Deployment could not roll out.
Highlights
Changes in this release:
- The controller can publish the gateway Vector config again — its own validation step could never complete on real clusters.
- The gateway Deployment now rolls out on clusters with as many system nodes as replicas, instead of stalling on a Pending pod.
- A destination whose credential cannot be resolved no longer takes the whole gateway config down with it.
- Gateway reloader metrics are renamed to stop colliding with log-shipper’s.
Improvements
This release improves:
- New
controller.vectorValidateTimeoutparameter (default60s) for thevector validaterun the controller performs before publishing a gateway config. - New alerts:
D8SecurityEventsManagerControllerConfigNotPublished,D8SecurityEventsManagerDestinationCredentialMissingandD8SecurityEventsManagerControllerValidateUnavailable. Previously the controller could stop publishing configs indefinitely without any alert firing.
Fixes
This release fixes:
- The controller no longer freezes the published gateway config. Its
vector validatestep ran with a 10s timeout in a 64Mi container and was either killed by the deadline or OOM-killed, and the failure was treated as a verdict that the config was invalid, so the Secret stayed at its last revision permanently. Validation now runs with--no-environment(~2s instead of ~5s on a 256 KB config, no topology build and no network healthchecks), the container floor is 192Mi, the timeout default is 60s, and a check that cannot be executed no longer blocks publishing — the gateway reloader validates the config regardless. - The gateway Deployment now uses
maxSurge: 0/maxUnavailable: 1. With the default surge, updating a 2-replica gateway asked for a third pod, and the required host anti-affinity meant that pod stayed Pending forever on a cluster with two system nodes: the module reported itself Deployed while still running the previous image. - A ClusterSecurityEventDestination that declares a credential the controller cannot resolve now blocks publishing and reports the exact destination and field, instead of emitting an
authblock without its token. Vector rejects such a block outright, which made the whole gateway config fail to load and stopped delivery to every destination, not just the affected one. Covers Loki, Elasticsearch, Splunk HEC and Kafka SASL. auth.strategy: None, the CRD default, no longer reaches the generated config as{"strategy": "none"}, which Vector does not accept.- The controller Deployment now has a
systemnode selector; it previously had none and could be scheduled onto worker nodes. - The CI Vector validation test now validates the
defaults.jsonthe gateway image actually ships instead of a hand-written stand-in.
Breaking changes
Changes that affect backward compatibility:
- Gateway reloader metrics are renamed with a
security_events_manager_gateway_prefix:vector_config_validation_error,vector_config_apply_failures_total,vector_config_periodic_retry_attempts_total,vector_config_reloads_success_total,vector_config_last_reload_success_timestamp_seconds,reloader_validate_duration_seconds,reloader_validate_timeout_seconds,reloader_apply_duration_seconds,reloader_vector_restarts_total,reloader_fsnotify_events_total. The old names collided with log-shipper’s vector-reloader, whose alerting rules do not filter by namespace, so this module’s config problems firedD8LogShipperConfigInvalid. Module rules, dashboards and docs are updated; custom recording rules, alerts and panels built on the old names need updating. - The
CONTROLLER_VECTOR_VALIDATE_SKIP_ABOVE_BYTESenvironment variable is removed. Its 256 KiB threshold sat just above the real config size, so it was liable to skip validation by accident rather than by intent; the new fail-open behaviour covers the case it was meant to. - The default
minimalSeveritychanged fromMediumtoLowin v1.0.3 without a release note. Clusters upgrading from v1.0.0 or earlier ship substantially more events than before, which interacts with the disk buffer defaults also introduced in v1.0.3. Set it explicitly if the previous volume is wanted.
Upgrade notes
Before upgrading, note the following:
- The controller container memory floor rises from 64Mi to 192Mi (VPA ceiling from 256Mi to 512Mi). This is what
vector validateneeds to run at all; below it the container is OOM-killed mid-validation.
v1.0.3
Release date: 2026-08-17
Gateway buffer configuration, enrichment plugins, SecretRefs for destinations, CEF/Socket sinks, and security hardening for the Vector gateway and controller RBAC.
Highlights
Changes in this release:
- Configurable Disk buffers for gateway and log-shipper layers, with per-destination overrides.
- Enrichment plugins (
k8s-pod-info,k8s-container-info,k8s-nodeuser-info) viaClusterSecurityEventEnrichmentPlugin. - Credentials for destinations can use SecretRefs instead of plaintext in the CR.
- CEF encoding and Socket destinations for SIEM integration.
- Gateway API bound to loopback; tighter Secret RBAC and SSRF checks on enrichment endpoints.
New features
This release adds:
- Configurable buffer settings:
gateway.bufferandgateway.logShipperBufferin module values; per-destination override viaspec.bufferin ClusterSecurityEventDestination; default Disk + Block;maxSizerequires a Kubernetes quantity with a unit suffix (for example512Mi). - Event
K8S_ADMISSION_POLICY_DENIEDfor admission-policy-engine denials in the Kubernetes audit log. - CRD
ClusterSecurityEventEnrichmentPluginfor built-in Internal enrichment plugins served by theenrichment-cachesidecar (External plugins are in the schema but not wired yet). - Enrichment plugins
k8s-pod-info,k8s-container-info, andk8s-nodeuser-info. tokenSecretRefandpasswordSecretRefon ClusterSecurityEventDestination for Loki, Elasticsearch, Kafka SASL, and Splunk HEC; built-incluster-lokistores its token in a Secret.- CEF encoding for Kafka, Vector, File, and Console destinations (
encoding.codec: CEF), optional syslog wrapping, and default CEF metadata viaspec.cef/ module values. - Socket destination type (TCP/UDP/Unix) for CEF-over-syslog to SIEM systems.
Fixes
This release fixes:
- Fixed the
D8SecurityEventsManagerControllerReconcileErrorsHighalert. gateway.reloaderValidateTimeoutconfigures the Vector validate timeout (default 120s) so large gateway configs do not fail with a silent timeout.- TLS Secrets in
d8-log-shipperare cleaned up when the related SecurityEventDefinition is deleted. - Credential Secret resolution failures are logged at warning with field path context.
- Credential Secret watcher uses label
security-events-manager.deckhouse.io/credential-secret: trueinstead of watching all Secrets. - Reloader
SaveTo()writes the Vector config with mode0600. resolveDestinationSecretRefs()returns a deep copy so credentials stay fresh after Secret rotation.- Lua
extractJSONreturns nil and logs whencjsonis unavailable (no fragile regex fallback). - Corrected the enrichment-cache port comment (9260 → 9261).
- E2E triggers for
K8S_CLUSTERROLEBINDING_DELETEDandK8S_CONFIGMAP_MODIFIEDcreate the test namespace and guard CRB delete against 404 races. - Controller-created Secrets and log-shipper CRs include the
heritage: deckhouselabel.
Security updates
Security updates in this release:
- Fixed CVEs in module images.
- Vector gateway API listens on
127.0.0.1:8686only; port 8686 removed from the Service. - mTLS client private key is stored in a Secret in
d8-log-shipper, not embedded in ClusterLogDestination. - Controller Secret RBAC limited to
d8-security-events-managerandd8-log-shipper. - CEL blocks SSRF on ClusterSecurityEventEnrichmentPlugin
spec.endpoint.url. gateway-vector-configSecret annotated withsecurity.deckhouse.io/contains-credentials: true.
Breaking changes
Changes that affect backward compatibility:
- Shipper
enrichshape:pluginis a free-form string;argsis an ordered{key,value}array; plugin names use hyphens (k8s-pod-info). Old CRs with the previous shape fail to parse; Plugin enrichment was not functional before. spec.splunkHEC.tokenis optional whentokenSecretRefis set.- CEL enforces mutual exclusivity of inline
token/passwordand their*SecretRefcounterparts on ClusterSecurityEventDestination.
v1.0.0
Release date: 2026-06-26
Initial release of the security-events-manager module.
Highlights
Changes in this release:
- First public release of the module.