The module lifecycle stage: Experimental

The module has requirements for installation

v1.2.0

Release date: 2026-09-28

The release reduces the log collection load on nodes, drops security events that fail the event registry check, and restricts the module schemas and permissions. The module requires Deckhouse 1.76 and log-shipper 1.0.6 or later.

Highlights

Changes in this release:

  • Each log source is read once per node regardless of the number of event codes.
  • Events that fail the event registry check are dropped, and the enrichment plugins fill their target fields.
  • The module custom resources and configuration are bounded by the schema, the component permissions are narrowed, and destination fields that can contain secrets are hidden from users without access to them.
  • A cluster-wide shipper collects pod logs from the namespaces it specifies, and a namespace selector can include and exclude namespaces at the same time.
  • A destination that sends events in the CEF format can specify the format version in the header prefix.

New features

This release adds:

  • The D8SecurityEventsManagerAgentDeliveryStalled alert fires when a log-shipper agent has undelivered security events and has not delivered any of them to the gateway for 10 minutes. Previously, such a delay triggered no alert, and the gateway pods remained Ready. The alert specifies the node and the sink.
  • The input.kubernetesPods.namespaceSelector parameter accepts matchNames and excludeNames together: the exclusion is applied after the inclusion. Previously, only one of the lists could be specified.
  • The encoding.cef.version parameter of ClusterSecurityEventDestination sets the CEF format version: V0 writes the CEF:0 prefix and is used by default, V1 writes CEF:1. The header fields and extensions do not change. The parameter is available for the Kafka, Vector, File, Console, and Socket destinations.
  • The spec.cef.version parameter of ClusterSecurityEventConfig sets the CEF format version for the whole cluster. The encoding.cef.version parameter of a destination overrides it.
  • In DKP 1.78 and later, with the multitenancy-manager module enabled, SecurityEventDefinition and ClusterSecurityEventEnrichmentPlugin are registered as grantable cluster resources. By default, both resources are available to all projects. To restrict access to them, use ClusterResourceGrantPolicy.

Improvements

This release improves:

  • Collection configs and destinations are created per shipper item instead of per event code. Previously, a log source with 50 event codes was read 50 times on every node and required a separate disk buffer for each code, up to about 20 GiB per node. Events are still filtered on the node.
  • All destinations use one TLS Secret in the d8-log-shipper namespace instead of a copy per event code.
  • The gateway determines the event code by the same extract rules as the agent. A log line that matches several event codes produces an event for each of them.
  • If a shipper item contains a rule that cannot be converted into a node filter, the logs of the item are sent to the gateway unfiltered and are processed there. This prevents events from being dropped on the node.
  • Collection configs and destinations created by earlier versions of the module are removed after the new ones are ready.

Fixes

This release fixes:

  • Connections of log-shipper agents to the gateway no longer fail with the SSL alert number 80 error. The error occurred on about every second reconnection and delayed the delivery of events.
  • The k8s-pod-info, k8s-container-info, and k8s-nodeuser-info enrichment plugins fill the target field again. Starting from version 1.0.2, the plugins left the field empty, and the gateway log contained the cjson unavailable error: actor.id, object.name, and object.namespace of events from external modules and static user names were not set. If the JSON library fails to load, the gateway log now contains the reason.
  • Setting one encoding.cef parameter in ClusterSecurityEventDestination no longer sets the other parameters to built-in values. Previously, deviceVendor, deviceProduct, and deviceVersion received default values and overrode spec.cef of ClusterSecurityEventConfig.
  • If several SecurityEventDefinition objects declare the same event code and source, the gateway uses the definition with the alphabetically first name. Previously, the severity, category, and description of such events could change between reconciliations. The other definitions are reported by the gateway_config_problematic_cluster_objects metric with the duplicate_code_source problem.
  • The gateway configuration no longer changes between reconciliations when the cluster has several ClusterSecurityEventConfig objects. Previously, the order of route conditions and the source of the CEF header defaults depended on the order of objects in the API cache, which caused unnecessary gateway reloads and could change the vendor, product, and version in the CEF header.
  • The order of parsing rules in the gateway no longer changes between reconciliations. Previously, when two rules covered the same pods or two shippers declared the same source and event code, the applied parser or transform could change.
  • The cluster-loki destination is created only when the kube-rbac-proxy CA is available. Previously, the destination could be created without a CA, and all events sent to Loki were lost.
  • parser rules are applied to a shipper that selects pods by input.kubernetesPods.labelSelector.matchExpressions. Previously, the logs of such a shipper were collected but not parsed, and the fields set in transform remained empty.
  • A shipper with a pod selector that contains only NotIn and DoesNotExist requirements covers pods without labels, as Kubernetes and log-shipper do. Previously, the logs of such pods were collected but not parsed.
  • A ClusterSecurityEventShipper with input.type: KubernetesPods and namespaces specified in input.kubernetesPods.namespace or input.kubernetesPods.namespaceSelector.matchNames collects events from these namespaces. Previously, such a shipper remained in the ReconcileFailed state and collected no events.
  • The controller no longer logs the failed to remove finalizer during SED cleanup and failed to update success status errors when two reconciliations process the same SecurityEventDefinition at the same time. These errors did not affect the result of the reconciliation.
  • The gateway.logShipperBuffer.maxEvents parameter is applied. Previously, it was ignored, and a buffer of the Memory type used the log-shipper default.
  • TLS Secrets created in d8-log-shipper by earlier versions of the module are removed when no destination references them. When a SecurityEventDefinition is deleted, its destination is removed before the Secrets it references. Previously, a reference to a deleted Secret could break the log-shipper configuration for all modules in the cluster.
  • Changing or removing a shipper item or deleting a shipper removes the related ClusterLogDestination. Previously, such objects remained in the cluster, were processed by every log-shipper agent, and prevented the shared TLS Secret from being removed.
  • A shipper with a source that has no SecurityEventDefinition no longer blocks the removal of its other objects. Previously, items removed from such a shipper kept their collection configs and continued to collect logs.
  • The gateway drops events that fail the event registry check. Previously, with the default Low severity threshold, events without a matching SecurityEventDefinition or without a required field reached the storage without a severity, category, and description, and severity-based filters and dashboards did not show them. The D8SecurityEventsManagerGatewayEventRegistryDrops alert fires on such drops.

Security updates

Security updates in this release:

  • Destination fields that can contain secrets are marked with x-kubernetes-sensitive-data: true: the authentication token and password of the http, loki, and elasticsearch destinations, the Kafka SASL password, the Splunk HEC token, and the spec.http.headers values. In Deckhouse 1.76 and later, these fields are removed from get, list, and watch responses for users without access to the clustersecurityeventdestinations/sensitive subresource and are masked in the audit log. In Deckhouse 1.77 and later, they are also masked in create, update, and patch responses. The values are stored in etcd unencrypted unless controlPlaneManager.encryptionEnabled is enabled, so use the tokenSecretRef and passwordSecretRef fields instead.
  • The bundled cluster-loki destination verifies the certificate and the host name of Loki. Previously, in Deckhouse versions earlier than 1.76, the verification was disabled, which allowed Loki to be impersonated from the cluster network.
  • All user-settable fields of the module custom resources and configuration have explicit limits: maxLength for strings, maxItems for arrays, maxProperties for maps, and minimum and maximum for numbers.
  • SecurityEventDefinition.spec.code and produces[].eventCode of both shipper types must match the ^[A-Z0-9_]+$ format. Previously, objects with other codes were accepted and ignored. Such objects are now rejected on creation and on a spec change.
  • The controller ClusterRole no longer grants cluster-wide access to secrets and events. Access to Secrets is limited to the module namespace and d8-log-shipper, and write access in the module namespace is limited to gateway-vector-config and gateway-destination-cas. The controller no longer changes Secrets owned by users.
  • The gateway reloader and the enrichment-cache sidecar serve metrics only on loopback addresses: 127.0.0.1:9256 and 127.0.0.1:9261. Ports 9255 and 9260, which are reachable from the cluster network, serve only health probes. Prometheus collects the metrics through kube-rbac-proxy.
  • The controller Service no longer publishes port 8081 of the health probes.
  • RBAC objects are renamed to include the component name: d8:security-events-manager:controller:rbac-proxy, d8:security-events-manager:aggregator:rbac-proxy, d8:security-events-manager:gateway:rbac-proxy, d8:security-events-manager:gateway:pod-reader, and d8:security-events-manager:controller:log-shipper-secrets for the Role and RoleBinding in d8-log-shipper.

Breaking changes

Changes that affect backward compatibility:

  • The module requires log-shipper 1.0.6 or later.
  • The minimum supported Deckhouse version is 1.76. The module is not installed on earlier versions.
  • A namespace name in input.kubernetesPods.namespaceSelector and input.kubernetesPods.namespace must be a valid DNS-1123 label. Previously, invalid names were accepted and broke log collection for the entire shipper.
  • input.kubernetesPods.labelSelector.matchLabels cannot be empty. Previously, an empty matchLabels was accepted, and logs were collected from all namespaces without producing events.
  • A requirement in input.kubernetesPods.labelSelector.matchExpressions must have a non-empty key, must have values for the In and NotIn operators, and must not have values for the Exists and DoesNotExist operators. Previously, such requirements were accepted, and the shipper remained in the ReconcileFailed state.
  • The module configuration is bounded by the schema: replicas has minimum and maximum values, and CPU and memory values must use the plain or suffixed Kubernetes quantity format. The exponential format, for example 128e6, is rejected. A ModuleConfig that does not match the schema blocks the module update.
  • SecurityEventDefinition.spec.metadata accepts only docs.desc.{en,ru} and tests.{executionTarget,testability,trigger.commands,trigger.cleanupCommands}. Other keys are removed on the next write of the object.

Upgrade notes

Before upgrading, note the following:

  • Update Deckhouse to 1.76 or later before updating the module.
  • Update log-shipper to 1.0.6 or later before updating the module. With an earlier log-shipper version, the module keeps the previous collection scheme.
  • Before updating, check the ModuleConfig, your SecurityEventDefinition objects with other keys in spec.metadata, and custom resources with long regular expressions, Grok patterns, field paths, or HTTP header values against the new schemas. The bundled destination and event definitions match the new schemas.
  • Before updating, check shippers for namespace names and matchExpressions requirements that the new schema rejects. Existing objects with such values are not deleted, but any change to them is rejected until the values are corrected.
  • Find SecurityEventDefinition objects and shippers with event codes that do not match ^[A-Z0-9_]+$. They are ignored, and any change to their spec is rejected.
  • To read inline destination credentials, grant get, list, or watch access to the clustersecurityeventdestinations/sensitive subresource. The module controller and the d8:security-events-manager:admin-kubeconfig role already have this access.
  • Add the security-events-manager.deckhouse.io/credential-secret label to manually created Secrets with credentials of alert rule destinations. The controller no longer adds this label. The Secret bundled with the module already has it.
  • To make destinations created before the update follow spec.cef of ClusterSecurityEventConfig, remove deviceVendor, deviceProduct, and deviceVersion from their encoding.cef. The values written by earlier versions remain in the objects.
  • After the update, expect a one-time duplication of events from the collected log files: each file is read again from the beginning of its current rotation.
  • After the update, events that fail the event registry check do not reach the storage. If the D8SecurityEventsManagerGatewayEventRegistryDrops alert fires, find the dropped event code by the enforce_event_registry: dropped line in the log of the vector container of the gateway, and make sure that the source fills the fields that the SecurityEventDefinition lists as required.

Docs

Documentation changes:

  • The examples page contains a ClusterSecurityEventShipper that collects pod logs from specified namespaces, including matchNames and excludeNames used together.
  • The SecurityEventLoggingTransformationRules example is placed in the d8-runtime-audit-engine namespace, where the shipper that references it runs. The page also states that spec.selector of these rules is not used to select pods.
  • The k8s-pod-info examples use @root.pod and @root.namespace for the pod name and namespace. The previous examples used fields that events do not contain.
  • The Falco examples select pods by the app: runtime-audit-engine label in the d8-runtime-audit-engine namespace. The previous examples used a label that module pods do not have.
  • The path syntax of transform[].value and enrich[].args[].value is described, including brackets and quotes for keys that contain a dot. The Falco examples are corrected accordingly.

v1.1.5

Release date: 2026-09-07

Security updates for the alert-evaluator, controller, and go-hooks images: dependency upgrades remediate known vulnerabilities, and OpenVEX attestations mark unused golang.org/x/crypto packages as not affected.

Highlights

Changes in this release:

  • The alert-evaluator, controller, and go-hooks images include updated golang.org/x libraries that remediate High and Critical vulnerabilities found in recent image scans.
  • OpenVEX attestations for the controller and module bundle mark vulnerabilities in golang.org/x/crypto/ssh and golang.org/x/crypto/openpgp as not affected, because those packages are not used.

Security updates

Security updates in this release:

  • In the alert-evaluator image, golang.org/x/net is updated to v0.56.0, golang.org/x/text to v0.39.0, and golang.org/x/sys to v0.46.0. The update remediates related CVEs, including CVE-2025-47911, CVE-2025-58190, CVE-2026-25680, CVE-2026-25681, CVE-2026-27136, CVE-2026-33814, CVE-2026-39821, CVE-2026-39824, CVE-2026-42502, CVE-2026-42506, CVE-2026-46600, and CVE-2026-56852.
  • In the controller and go-hooks images, golang.org/x/crypto is updated to v0.55.0. The update remediates Critical CVE-2026-56854.
  • OpenVEX statements for the controller and module bundle mark CVE-2026-56855, CVE-2026-78662, and GO-2026-5932 as not_affected with the vulnerable_code_not_in_execute_path justification. The controller uses golang.org/x/crypto for bcrypt, scrypt, and pbkdf2 through sprig; go-hooks uses it for ocsp and pkcs12 through cfssl. Neither component imports golang.org/x/crypto/ssh or golang.org/x/crypto/openpgp.

v1.1.4

Release date: 2026-09-03

Alert rules are now LogQL, evaluated directly against Loki: ClusterSecurityEventAlertRule is reshaped after PrometheusRule (groups of rules with alert/expr/for/labels/annotations), and the alerting pipeline no longer runs through Vector.

Highlights

Changes in this release:

  • ClusterSecurityEventAlertRule is now shaped like PrometheusRule: spec.groups[].rules[], each rule with alert, for, severityLevel, labels and annotations, and either a hand-written LogQL expr or a structured match/aggregation block (event codes, sources, categories, actors, grouping, threshold) that compiles into the equivalent LogQL. This is a breaking change to the resource spec — rewrite existing rules before upgrading.
  • Every rule declares its form in the required mode field (Expr or Match), matching the block it sets. The platform web interface uses mode to show only the fields for the selected form.
  • A new alert-evaluator component, running inside the existing aggregator Deployment, evaluates each rule’s LogQL query against the cluster’s Loki on a schedule and reports firing rules through the same security_events_alert_last_firing_timestamp_seconds metric as before. The D8SecurityEventAlertFiring alert is unchanged.
  • gateway no longer forwards a copy of matching events to aggregator over the network: alert-evaluator reads directly from Loki. aggregator’s TLS material and Vector process are no longer needed.
  • Alert rules require the cluster-loki ClusterSecurityEventDestination to be enabled — without it, LogQL has no data to query.

Improvements

This release improves:

  • A rule can now look arbitrarily far back and group by any label LogQL supports, not only the eight fixed fields the previous aggregation.groupBy enum offered. Use a hand-written expr for that.
  • Each ClusterSecurityEventAlertRule reports its own status.conditions (Ready/Degraded, with a reason such as LogQLQueryError, LokiDestinationMissing or CardinalityLimitExceeded).

Breaking changes

Changes that affect backward compatibility:

  • Every rule in ClusterSecurityEventAlertRule.spec.groups[].rules[] must declare mode: Expr or mode: Match, matching the block it sets (expr or match). A rule without mode, or with a mode that disagrees with its block, is rejected.
  • The old spec.match/spec.aggregation/spec.alert fields (one rule per object) are removed. Rewrite each rule under spec.groups[].rules[], using either a hand-written LogQL expr or the same match/aggregation fields as before (eventCodes, sources, categories, severityMin, actorType, actors, excludeActors, excludeSystemActors, groupBy, threshold, window), now nested per rule. See the module documentation for examples.
  • The API server now enforces field constraints on ClusterSecurityEventShipper, PodSecurityEventShipper, ClusterSecurityEventLogTransformation, SecurityEventLogTransformation, SecurityEventDefinition, ClusterSecurityEventEnrichmentPlugin and ClusterSecurityEventDestination that were previously silent. An empty label selector (labelSelector: {} / selector: {}) is now rejected instead of matching everything, ClusterSecurityEventDestination’s spec.*.buffer.maxEvents must be at least 100, and several free-text fields no longer accept an empty string.
  • Alert rule evaluation requires the cluster-loki ClusterSecurityEventDestination (enable it via securityEventConfig.destinations). Without it, every ClusterSecurityEventAlertRule is reported Degraded and no rule is evaluated; event collection and delivery to other destinations are unaffected.

Upgrade notes

Before upgrading, note the following:

  • When rewriting a rule, add mode: Expr to a rule with a hand-written expr, or mode: Match to a rule built from match/aggregation.
  • Before upgrading, save every existing ClusterSecurityEventAlertRule (d8 k get clustersecurityeventalertrule -o yaml). After upgrading, rewrite each rule under the new spec.groups[].rules[] shape — the old top-level spec fields are rejected by the updated CRD schema.
  • Before upgrading, check whether existing resources violate the newly enforced constraints: run d8 k get clustersecurityeventshipper,podsecurityeventshipper,clustersecurityeventlogtransformation,securityeventlogtransformation,clustersecurityeventdestination -o yaml and look for an empty labelSelector/selector or buffer.maxEvents below 100. An existing object that violates a constraint keeps working, but the next write to it is rejected.
  • Confirm cluster-loki is enabled (securityEventsManager.securityEventConfig.destinations includes cluster-loki) before relying on alert rules; otherwise every rule stays Degraded.

Known issues

Known limitations of this release:

  • Evaluation state (pending/firing) lives in the evaluator process’s memory: restarting it resets the for countdown for every rule.
  • Grouping by a LogQL label with unbounded cardinality is capped at 500 distinct label sets per rule; the rest are dropped and the rule is reported CardinalityLimitExceeded, but it keeps firing for the groups it already tracks.

Docs

Documentation changes:

  • The alerting section of the examples page is rewritten for the PrometheusRule-shaped resource, covering both the expr and match/aggregation forms, including the line filter that reproduces the old excludeSystemActors behavior.

v1.1.3

Release date: 2026-08-27

Security fixes for module images: OpenVEX attestations mark golang.org/x/crypto GO-2026-5932 as not affected for the controller and module bundle hooks.

Highlights

Changes in this release:

  • Module images publish OpenVEX attestations so CVE scans suppress GO-2026-5932 where openpgp is not on the execute path.

Security updates

Security updates in this release:

  • Added OpenVEX statements for GO-2026-5932 (golang.org/x/crypto/openpgp) on the controller and module bundle (go-hooks), status not_affected / vulnerable_code_not_in_execute_path.
  • Wired cosign OpenVEX attestation into the module build so CI CVE scans read the statements from the OCI registry.

v1.1.2

Release date: 2026-08-27

Alerting on security events: the new ClusterSecurityEventAlertRule resource counts matching events per key over a time window and raises an ordinary Prometheus alert once the threshold is reached. Adds an Http destination type for webhooks and HTTP collectors, and a new event for containerd integrity violations.

Highlights

Changes in this release:

  • New ClusterSecurityEventAlertRule resource: describe which events to watch, how to count them, and what the alert should say.
  • A firing rule raises the Prometheus alert D8SecurityEventAlertFiring, so it becomes a ClusterAlert object, shows up in the console next to every other alert, keeps its history and travels through the cluster’s existing alert delivery. There is nothing to configure for delivery.
  • New Http destination type sends security events to any HTTP endpoint — a webhook or an HTTP collector — with authentication, TLS, custom headers and a choice of body framing.
  • New security event D8_CONTAINERD_INTEGRITY_VIOLATION reports a containerd integrity check failure — most often an image refused at pull because its signature did not verify.

New features

This release adds:

  • ClusterSecurityEventAlertRule (cluster-scoped, short name csear). spec.match selects events by event code, minimum severity, category, source component, actor and actor type; spec.aggregation groups them by groupBy and fires as soon as a group reaches threshold within window; spec.alert carries the summary, description, extra labels and the DKP severityLevel. A rule without aggregation alerts on every matched event.
  • match.excludeSystemActors (enabled by default) keeps platform identities out of a rule, using the same list as the “Exclude system” filter of the console. Routine cluster work dominates the event stream — on an idle cluster the DKP ServiceAccount alone produced 2617 of 2812 secret reads in an hour — so without it the first rule anyone writes is an alert about the platform. match.excludeActors and match.actors take globs against actor.id for anything the built-in list does not cover.
  • The alert resolves on its own about five minutes after the rule last fired. There is no field to set: the aggregator publishes security_events_alert_last_firing_timestamp_seconds, the moment a rule last fired, and D8SecurityEventAlertFiring tests the age of it. The metric is exported on the aggregator’s /alert-metrics/metrics endpoint and can be used in dashboards.
  • New destination type Http: spec.http.endpoint with optional auth, tls and headers, and framing — JSONArray (default) for receivers that expect one document per request, NewlineDelimited for collectors that read NDJSON. The number of delivery attempts for one batch is bounded, so a receiver that keeps rejecting a request it will never accept cannot hold up everything queued behind it: once the attempts run out the batch is dropped into the sink’s error metrics, which the module already alerts on.
  • New security event D8_CONTAINERD_INTEGRITY_VIOLATION (category Runtime, severity High, source containerd-integrity). It fires when containerd refuses an image whose signature does not verify, and when a file checksum on the node does not match the expected one. object.name carries the image reference, object.namespace the namespace it was being pulled into, and metadata.extra the message, the file path and the expected and actual hashes where the event provides them. Collected only on nodes running cri ContainerdV2 in the CSE edition; probe traffic from d8-upmeter is filtered out.
  • kube-audit events now carry actor.sourceIP, the originating client address taken from the head of the request chain, next to the full actor.sourceIPs chain. Alert rules can group by it (ActorSourceIP), which is how a rule about unauthorized requests identifies the client when there is no identity to group on.
  • The module runs a new single-replica alert-aggregator component that holds the counters, with its own alert D8SecurityEventsManagerAggregatorNotAvailable. While it is down, event collection and delivery continue — events wait in the gateway buffer — but counting windows restart, so a partially reached threshold begins from zero.

Improvements

This release improves:

  • Transform source paths can address array elements, for example sourceIPs[0]. Previously every path segment was treated as a map key, so any array field in a log was unmappable.
  • Module documentation, CRDs and OpenAPI schemas are now included in the bundle and release images.

Upgrade notes

Before upgrading, note the following:

  • No configuration change is required. The alerting components are created on upgrade and stay idle until the first ClusterSecurityEventAlertRule exists. The alert-aggregator runs a single replica with a Recreate rollout by design — counters live in one process — so its updates and node drains produce a short gap in counting.

Known issues

Known limitations of this release:

  • A rule can only match what the cluster’s audit policy records. The default DKP policy logs list of secrets but not a get of one secret by name, and drops most requests from system:authenticated, so a rule on K8S_SECRET_ACCESSED will not see d8 k get secret <name>. Check the audit policy before concluding that a rule does not work.
  • Counters live in the aggregator’s memory. Restarting it starts every window from zero.
  • Grouping by a value with unbounded cardinality is capped: past 500 distinct values for a label the tag is dropped and the alert still fires, one dimension poorer.
  • On clusters with update.blockOnAlerts enabled in the deckhouse module, an alert at or below the configured level (4 by default) blocks DKP releases from being applied while it is firing. The severityLevel default of 6 stays out of the way; lower it deliberately.

Docs

Documentation changes:

  • The examples page gained an alerting section with a rule per detection story: repeated exec into pods by one actor, a burst of RBAC edits grouped per actor and namespace, unauthorized requests grouped by client address, and a containerd integrity violation as the case where a single occurrence is enough. It also gained an HTTP destination example.

v1.0.5

Release date: 2026-08-24

Security fixes for module images: updated Go dependencies in batch hooks and related components to remediate known CVEs.

Highlights

Changes in this release:

  • Module images include dependency updates that address known CVEs in Go components.

Security updates

Security updates in this release:

  • Bumped golang.org/x/net and golang.org/x/text in hooks/batch and related Go images for CVE remediation.
  • CVE-oriented security fixes included in the component images.

v1.0.4

Release date: 2026-08-18

Fixes an upgrade deadlock: the controller could not publish gateway configs and the gateway Deployment could not roll out.

Highlights

Changes in this release:

  • The controller can publish the gateway Vector config again — its own validation step could never complete on real clusters.
  • The gateway Deployment now rolls out on clusters with as many system nodes as replicas, instead of stalling on a Pending pod.
  • A destination whose credential cannot be resolved no longer takes the whole gateway config down with it.
  • Gateway reloader metrics are renamed to stop colliding with log-shipper’s.

Improvements

This release improves:

  • New controller.vectorValidateTimeout parameter (default 60s) for the vector validate run the controller performs before publishing a gateway config.
  • New alerts: D8SecurityEventsManagerControllerConfigNotPublished, D8SecurityEventsManagerDestinationCredentialMissing and D8SecurityEventsManagerControllerValidateUnavailable. Previously the controller could stop publishing configs indefinitely without any alert firing.

Fixes

This release fixes:

  • The controller no longer freezes the published gateway config. Its vector validate step ran with a 10s timeout in a 64Mi container and was either killed by the deadline or OOM-killed, and the failure was treated as a verdict that the config was invalid, so the Secret stayed at its last revision permanently. Validation now runs with --no-environment (~2s instead of ~5s on a 256 KB config, no topology build and no network healthchecks), the container floor is 192Mi, the timeout default is 60s, and a check that cannot be executed no longer blocks publishing — the gateway reloader validates the config regardless.
  • The gateway Deployment now uses maxSurge: 0 / maxUnavailable: 1. With the default surge, updating a 2-replica gateway asked for a third pod, and the required host anti-affinity meant that pod stayed Pending forever on a cluster with two system nodes: the module reported itself Deployed while still running the previous image.
  • A ClusterSecurityEventDestination that declares a credential the controller cannot resolve now blocks publishing and reports the exact destination and field, instead of emitting an auth block without its token. Vector rejects such a block outright, which made the whole gateway config fail to load and stopped delivery to every destination, not just the affected one. Covers Loki, Elasticsearch, Splunk HEC and Kafka SASL.
  • auth.strategy: None, the CRD default, no longer reaches the generated config as {"strategy": "none"}, which Vector does not accept.
  • The controller Deployment now has a system node selector; it previously had none and could be scheduled onto worker nodes.
  • The CI Vector validation test now validates the defaults.json the gateway image actually ships instead of a hand-written stand-in.

Breaking changes

Changes that affect backward compatibility:

  • Gateway reloader metrics are renamed with a security_events_manager_gateway_ prefix: vector_config_validation_error, vector_config_apply_failures_total, vector_config_periodic_retry_attempts_total, vector_config_reloads_success_total, vector_config_last_reload_success_timestamp_seconds, reloader_validate_duration_seconds, reloader_validate_timeout_seconds, reloader_apply_duration_seconds, reloader_vector_restarts_total, reloader_fsnotify_events_total. The old names collided with log-shipper’s vector-reloader, whose alerting rules do not filter by namespace, so this module’s config problems fired D8LogShipperConfigInvalid. Module rules, dashboards and docs are updated; custom recording rules, alerts and panels built on the old names need updating.
  • The CONTROLLER_VECTOR_VALIDATE_SKIP_ABOVE_BYTES environment variable is removed. Its 256 KiB threshold sat just above the real config size, so it was liable to skip validation by accident rather than by intent; the new fail-open behaviour covers the case it was meant to.
  • The default minimalSeverity changed from Medium to Low in v1.0.3 without a release note. Clusters upgrading from v1.0.0 or earlier ship substantially more events than before, which interacts with the disk buffer defaults also introduced in v1.0.3. Set it explicitly if the previous volume is wanted.

Upgrade notes

Before upgrading, note the following:

  • The controller container memory floor rises from 64Mi to 192Mi (VPA ceiling from 256Mi to 512Mi). This is what vector validate needs to run at all; below it the container is OOM-killed mid-validation.

v1.0.3

Release date: 2026-08-17

Gateway buffer configuration, enrichment plugins, SecretRefs for destinations, CEF/Socket sinks, and security hardening for the Vector gateway and controller RBAC.

Highlights

Changes in this release:

  • Configurable Disk buffers for gateway and log-shipper layers, with per-destination overrides.
  • Enrichment plugins (k8s-pod-info, k8s-container-info, k8s-nodeuser-info) via ClusterSecurityEventEnrichmentPlugin.
  • Credentials for destinations can use SecretRefs instead of plaintext in the CR.
  • CEF encoding and Socket destinations for SIEM integration.
  • Gateway API bound to loopback; tighter Secret RBAC and SSRF checks on enrichment endpoints.

New features

This release adds:

  • Configurable buffer settings: gateway.buffer and gateway.logShipperBuffer in module values; per-destination override via spec.buffer in ClusterSecurityEventDestination; default Disk + Block; maxSize requires a Kubernetes quantity with a unit suffix (for example 512Mi).
  • Event K8S_ADMISSION_POLICY_DENIED for admission-policy-engine denials in the Kubernetes audit log.
  • CRD ClusterSecurityEventEnrichmentPlugin for built-in Internal enrichment plugins served by the enrichment-cache sidecar (External plugins are in the schema but not wired yet).
  • Enrichment plugins k8s-pod-info, k8s-container-info, and k8s-nodeuser-info.
  • tokenSecretRef and passwordSecretRef on ClusterSecurityEventDestination for Loki, Elasticsearch, Kafka SASL, and Splunk HEC; built-in cluster-loki stores its token in a Secret.
  • CEF encoding for Kafka, Vector, File, and Console destinations (encoding.codec: CEF), optional syslog wrapping, and default CEF metadata via spec.cef / module values.
  • Socket destination type (TCP/UDP/Unix) for CEF-over-syslog to SIEM systems.

Fixes

This release fixes:

  • Fixed the D8SecurityEventsManagerControllerReconcileErrorsHigh alert.
  • gateway.reloaderValidateTimeout configures the Vector validate timeout (default 120s) so large gateway configs do not fail with a silent timeout.
  • TLS Secrets in d8-log-shipper are cleaned up when the related SecurityEventDefinition is deleted.
  • Credential Secret resolution failures are logged at warning with field path context.
  • Credential Secret watcher uses label security-events-manager.deckhouse.io/credential-secret: true instead of watching all Secrets.
  • Reloader SaveTo() writes the Vector config with mode 0600.
  • resolveDestinationSecretRefs() returns a deep copy so credentials stay fresh after Secret rotation.
  • Lua extractJSON returns nil and logs when cjson is unavailable (no fragile regex fallback).
  • Corrected the enrichment-cache port comment (9260 → 9261).
  • E2E triggers for K8S_CLUSTERROLEBINDING_DELETED and K8S_CONFIGMAP_MODIFIED create the test namespace and guard CRB delete against 404 races.
  • Controller-created Secrets and log-shipper CRs include the heritage: deckhouse label.

Security updates

Security updates in this release:

  • Fixed CVEs in module images.
  • Vector gateway API listens on 127.0.0.1:8686 only; port 8686 removed from the Service.
  • mTLS client private key is stored in a Secret in d8-log-shipper, not embedded in ClusterLogDestination.
  • Controller Secret RBAC limited to d8-security-events-manager and d8-log-shipper.
  • CEL blocks SSRF on ClusterSecurityEventEnrichmentPlugin spec.endpoint.url.
  • gateway-vector-config Secret annotated with security.deckhouse.io/contains-credentials: true.

Breaking changes

Changes that affect backward compatibility:

  • Shipper enrich shape: plugin is a free-form string; args is an ordered {key,value} array; plugin names use hyphens (k8s-pod-info). Old CRs with the previous shape fail to parse; Plugin enrichment was not functional before.
  • spec.splunkHEC.token is optional when tokenSecretRef is set.
  • CEL enforces mutual exclusivity of inline token/password and their *SecretRef counterparts on ClusterSecurityEventDestination.

v1.0.0

Release date: 2026-06-26

Initial release of the security-events-manager module.

Highlights

Changes in this release:

  • First public release of the module.