The module lifecycle stage: Generally available version

The module has requirements for installation

Triggers

Triggers (alerting rules) define the conditions for creating alerts when metric values deviate from expected thresholds.

Triggers are defined in rule groups as items in the spec.rules array. If a rule contains the alert field, it is treated as a trigger and is used to create alerts.

Types of rule groups with triggers

Three types of rule groups can define triggers:

Rule group type Scope Who has access
System rule groups (ClusterObservabilityMetricsRulesGroup) Cluster level DKP administrators
Project rule groups (ObservabilityMetricsRulesGroup) Project (namespace) level Users of the corresponding project
Standard propagated rule groups (ClusterObservabilityPropagatedMetricsRulesGroup) Created at the cluster level and automatically available in all projects Users of all projects

Rule group types:

Additional labels for alerts shipped with DKP

The observability module lets you apply additional labels to alerting rules shipped with DKP. To do this, use the ClusterObservabilityAlertAdditionalLabels resource.

Additional labels are applied only to alerts created using the ClusterObservabilityMetricsRulesGroup and ClusterObservabilityPropagatedMetricsRulesGroup resources with the heritage: deckhouse label.

To add labels to custom alerts created with ObservabilityMetricsRulesGroup or ClusterObservabilityMetricsRulesGroup, use the spec.rules.labels field.

Configuration examples:

  • Adding a label to all cluster alerts:

    apiVersion: observability.deckhouse.io/v1alpha1
    kind: ClusterObservabilityAlertAdditionalLabels
    metadata:
      name: all-alerts
    spec:
      alertSelector:
        matchExpressions:
          - key: alertname
            operator: Exists
      additionalLabels:
        example-label-name: example-label-value
  • Adding the severity=Info label to alerts with severity_level of 7, 8, and 9:

    apiVersion: observability.deckhouse.io/v1alpha1
    kind: ClusterObservabilityAlertAdditionalLabels
    metadata:
      name: severity-info-low-priority
    spec:
      alertSelector:
        matchExpressions:
          - key: severity_level
            operator: In
            values: ["7", "8", "9"]
      additionalLabels:
        severity: Info
  • Adding the team=custom label to alerts with the specified names:

    apiVersion: observability.deckhouse.io/v1alpha1
    kind: ClusterObservabilityAlertAdditionalLabels
    metadata:
      name: deckhouse-team-routing
    spec:
      alertSelector:
        matchExpressions:
          - key: alertname
            operator: In
            values:
              - D8CNIMisconfigured
              - D8DeckhouseIsNotOnReleaseChannel
              - D8DeckhouseIsNotOnReleaseChannel
      additionalLabels:
        team: custom

Trigger groups

Trigger groups are used to logically organize monitoring rules and manage their parameters at the rule set level.

Groups are convenient for combining triggers related to the same component, service, or project, as well as for applying a shared evaluation interval to all rules in the group.

Notifications

The observability module provides mechanisms for configuring alert notification delivery and controlling access to notification channels at both the cluster and project levels.

The following delivery channels are supported:

  • Email
  • Telegram
  • Slack
  • Webhook
  • ExpressMessenger
  • Zabbix

Connection parameters depend on the channel type and are configured through the corresponding Kubernetes resource.

Protection of channel credentials

Credential fields are protected by the x-kubernetes-sensitive-data marker: without rights to the <resource>/sensitive subresource, the API server returns them as <omitted> in get/list/ watch and masks them in the audit log. /sensitive access is granted to anyone who can create/edit the channel itself; levels without write access do not get it.

Requires the CRDSensitiveData feature gate on kube-apiserver — enabled by default starting with DKP 1.77; on earlier versions the marker has no effect and credentials stay readable.

Types of notification channels

Three types of notification channels are supported:

Channel type Scope Who can create
System channels (ClusterObservabilityNotificationChannel) Cluster level DKP administrators
Project channels (ObservabilityNotificationChannel) Project (namespace) level Users of the corresponding project
Standard propagated channels (ClusterObservabilityPropagatedNotificationChannel) Created at the cluster level and automatically available in all projects DKP administrators

Channel types:

  • System channels (ClusterObservabilityNotificationChannel): Used for cluster-level notification delivery. Available in the Deckhouse web UI under “System” → “System management” → “Monitoring” → “Notification settings” → “Notification channels”.

  • Project channels (ObservabilityNotificationChannel): Allow configuring notification delivery within a specific project. Available in the corresponding project in the Deckhouse web UI under “Monitoring” → “Notification settings” → “Notification channels”.

  • Standard propagated channels (ClusterObservabilityPropagatedNotificationChannel): Created at the cluster level and automatically become available in all projects for notification delivery. Use the ClusterObservabilityPropagatedNotificationChannel resource or the d8 CLI utility to create them.

Webhook channel HTTP client configuration

For channels with spec.type: Webhook, you can additionally configure outbound HTTP client settings in spec.webhook.httpConfig.

There are 3 mutually exclusive authentication options:

  • basicAuth
  • authorization
  • oauth2

Required and optional fields

  • Required for a Webhook channel: spec.webhook.url
  • spec.webhook.httpConfig is optional
  • If oauth2 is used, oauth2.clientId and oauth2.tokenUrl are required

In addition to authentication, httpConfig supports:

  • transport options (enableHttp2)
  • proxy settings (proxyUrl, noProxy, proxyFromEnvironment, proxyConnectHeader)
  • TLS settings (tlsConfig)
  • custom headers (httpHeaders)

For OAuth2, there are two configuration levels:

  • httpConfig.proxy* and httpConfig.tlsConfig apply to webhook delivery requests
  • httpConfig.oauth2.proxy* and httpConfig.oauth2.tlsConfig apply to OAuth2 token requests

The following fields are deprecated and ignored for security reasons. They are still accepted by the CRD schema for backward compatibility, but they have no effect:

  • followRedirects — redirects are never followed, because a redirect makes Alertmanager fetch a location chosen by the notification target.
  • the file-based fields passwordFile, credentialsFile, clientSecretFile, tlsConfig.caFile, tlsConfig.certFile, tlsConfig.keyFile and httpHeaders.<name>.files — the files are no longer read from the Alertmanager container. Use the corresponding inline field (password, credentials, clientSecret) or httpHeaders.<name>.values / .secrets instead.

Example with basicAuth:

apiVersion: observability.deckhouse.io/v1alpha1
kind: ClusterObservabilityNotificationChannel
metadata:
  name: webhook-channel-basic-auth
spec:
  type: Webhook
  webhook:
    url: https://hooks.example/webhook
    httpConfig:
      basicAuth:
        username: notify-user
        password: notify-secret

Example with authorization:

apiVersion: observability.deckhouse.io/v1alpha1
kind: ClusterObservabilityNotificationChannel
metadata:
  name: webhook-channel-authorization
spec:
  type: Webhook
  webhook:
    url: https://hooks.example/webhook
    httpConfig:
      authorization:
        type: Bearer
        credentials: opaque-token

Example with oauth2:

apiVersion: observability.deckhouse.io/v1alpha1
kind: ClusterObservabilityNotificationChannel
metadata:
  name: webhook-channel-oauth2
spec:
  type: Webhook
  webhook:
    url: https://hooks.example/webhook
    httpConfig:
      oauth2:
        clientId: my-client
        clientSecret: my-secret
        tokenUrl: https://idp.example/token
        scopes:
          - read
          - write
        endpointParams:
          audience: myapp

Zabbix channel configuration

For channels with spec.type: Zabbix, Alertmanager pushes alerts to a Zabbix Server or Zabbix Proxy in push mode, using the Zabbix sender (trapper) protocol.

Required and optional fields

  • Required: spec.zabbix.server — the address of the Zabbix Server or Zabbix Proxy.
  • Required: spec.zabbix.clusterName — an arbitrary string that becomes the technical name of the Zabbix Host that owns the pushed items and the imported template. Pick any value that’s unique among every cluster/channel pointed at the same Zabbix instance (for example, your cluster’s public domain, if you have one, or any other identifier your team already uses for this cluster).
  • spec.zabbix.port — the sender protocol port (default 10051).
  • spec.zabbix.hostGroup — the name of the Zabbix host group the cluster’s Host is placed into. Defaults to Deckhouse if unset.
  • The sender connection to Zabbix is encrypted by default. spec.zabbix.tls optionally configures it further (caFile, certFile, keyFile, serverName, insecureSkipVerify — for a private CA or a self-signed certificate; a public CA needs no configuration). Zabbix’s PSK encryption mode is not supported. spec.zabbix.disableTLS: true opts out entirely, for a Zabbix Server/Proxy with no TLS on its trapper port.
  • spec.zabbix.api — optional; enables automatic import of the bundled template via the Zabbix API (see “Automatic import via the Zabbix API” below). If unset, only the manual-import ConfigMap is produced.

When an alert stops firing

  • A resolved alert is sent once more with the value 0, which clears its trigger immediately.
  • The next notification no longer lists the alert in the discovery data, so Zabbix marks its item as lost and deletes it after the discovery rule’s lifetime of 1 hour; the alert then disappears from Latest data too.
  • As a fallback, the trigger also stops firing once its item has received nothing for 5 minutes (for example when the alert was silenced and no resolved notification was sent).

Because the freshness check relies on alerts being re-sent, spec.notification.repeatInterval for a Zabbix channel is capped at 2m (a longer value is silently reduced); the default of 30s is fine. spec.alert.groupByLabels is ignored for Zabbix channels: all alerts matched by the policy are sent together in every notification, which is what the Zabbix discovery rules require. Which alerts are sent is still controlled by spec.alert.selector.

The d8alerts.sender.heartbeat item is kept fresh by the platform’s always-on DeadMansSwitch alert, routed to the channel like any other alert — through its ObservabilityNotificationPolicy selector. A narrow selector (e.g. severity="critical") needs an explicit alertname="DeadMansSwitch" match added to keep the heartbeat working. Zabbix uses nodata() on that item to detect if the path goes down.

The heartbeat item exists only in the template of a ClusterObservabilityNotificationChannel. A namespaced ObservabilityNotificationChannel and a ClusterObservabilityPropagatedNotificationChannel never receive DeadMansSwitch, so their templates have no heartbeat item.

Registering the template in Zabbix: manual or automatic

The bundled template (zabbix_export format) already includes the cluster’s Host and Host Group (spec.zabbix.clusterName/hostGroup), so importing it leaves nothing else to configure by hand. There are two ways to get it into Zabbix:

Manual import (always available)

The controller always renders the template into a ConfigMap for the channel, regardless of spec.zabbix.api. The name and namespace depend on the channel kind:

Channel kind ConfigMap name Namespace
ObservabilityNotificationChannel <channel-name>-zabbix-template same as the channel
ClusterObservabilityNotificationChannel cluster-<channel-name>-zabbix-template module namespace
ClusterObservabilityPropagatedNotificationChannel propagated-<channel-name>-zabbix-template module namespace

Extract it and import it into Zabbix under Data collection → Templates → Import, for example for a project channel named my-zabbix-channel in namespace my-namespace:

kubectl get configmap my-zabbix-channel-zabbix-template -n my-namespace \
  -o jsonpath='{.data.template\.yaml}' > d8alerts-sender.yaml
Automatic import via the Zabbix API (optional)

Setting spec.zabbix.api makes the controller import the exact same rendered template into Zabbix directly, via its JSON-RPC API (configuration.import), in addition to (not instead of) the ConfigMap above:

spec:
  type: Zabbix
  zabbix:
    server: zabbix.example.com
    clusterName: my-cluster
    hostGroup: Deckhouse
    api:
      url: https://zabbix.example.com/api_jsonrpc.php
      token: <zabbix-api-token>
  • spec.zabbix.api.url — address of the Zabbix JSON-RPC API; if the scheme is omitted, https:// is assumed. This is a different address than spec.zabbix.server: a Zabbix Proxy has no API of its own, so api.url must point at the central Zabbix Server. Requires Zabbix ≥6.4.
  • spec.zabbix.api.token — a Zabbix API token, stored in this resource in plain text. Create a dedicated token with a narrowly-scoped role (limited to the relevant template group and host group) rather than reusing a Super Admin account.
  • spec.zabbix.api.tlsConfig — optional TLS settings for the connection to the Zabbix API (mirrors spec.zabbix.tls, but the certificate/key paths must be available inside the observability-controller container instead of the alertmanager container).

If the import fails, check the observability-controller logs for failed to import zabbix template via api. Typical causes: the Zabbix API is unreachable, the token is invalid or lacks permission on the relevant template or host group, or the Zabbix version is lower than 6.4. The import is retried automatically on every reconcile (e.g. after editing the channel, or the next time the controller restarts), so a transient failure clears itself without any action.

Choosing the Zabbix server address

Prefer a stable DNS name over a literal IP for spec.zabbix.server and spec.zabbix.api.url — for in-cluster Zabbix, a Service DNS name rather than a node or pod IP. The admission webhook emits a non-blocking warning for literal IPs.

Notification policies

Notification policies define which channel should be used to deliver notifications for an alert (or a group of alerts).

Policy type Description How to configure
System notification policies Used to configure delivery rules for system alerts. System policies can use only system notification channels. Available in the Deckhouse web UI under “System” → “System management” → “Monitoring” → “Notification settings” → “Notification policies”. Use the ClusterObservabilityNotificationPolicy resource.
Project notification policies Used to configure delivery rules for project alerts. Project policies can use project or standard cluster channels, but not system notification channels. Available in the corresponding project under “Monitoring” → “Notification settings” → “Notification policies”. Use the ObservabilityNotificationPolicy resource.

Notification silencing

In situations where notifications are expected in advance (for example, during planned maintenance or testing), the observability module allows disabling notification delivery for alerts matching specified conditions.

Silence type Description How to configure
System notification silences Used to configure silencing rules for system alert delivery. Available in the Deckhouse web UI under “System” → “System management” → “Monitoring” → “Notification settings” → “Notification silencing”. Use the ClusterObservabilityNotificationSilence resource.
Project notification silences Used to configure silencing rules for project alert delivery. Available in the corresponding project under “Monitoring” → “Notification settings” → “Notification silencing”. Use the ObservabilityNotificationSilence resource.

Alerts

The observability module provides access control separation for cluster-level and project-level alerts and allows viewing the list of active and resolved alerts.

Active alerts are grouped by severity level:

  • critical (critical, S1–S3)
  • warning (warning, S4–S6)
  • informational (info, S7–S9)

When viewing an alert, the user can see general information, labels, annotations, and a graph.

Types of alerts

Two types of alerts are supported:

Alert type Scope Who has access
System alerts (ClusterObservabilityAlerts) Cluster level DKP administrators
Project alerts (ObservabilityAlerts) Project (namespace) level Users of the corresponding project

Alert types:

  • System alerts (ClusterObservabilityAlerts): Relate to DKP cluster components. The full list of active and resolved system alerts is available in the Deckhouse web UI under “System” → “System management” → “Monitoring” → “Active alerts”.

  • Project alerts (ObservabilityAlerts): Relate to resources of a specific project (namespace). The full list of active and resolved project alerts is available in the Deckhouse web UI in the corresponding project under “Monitoring” → “Active alerts”.

DeadMansSwitch and PrometheusUnavailable alerts

DeadMansSwitch

DeadMansSwitch is a service alert that fires continuously, confirming the normal operation of Prometheus and the entire alert delivery pipeline.

If DeadMansSwitch stops arriving, the PrometheusUnavailable alert starts firing.

By default, the DeadMansSwitch alert is sent to all configured notification channels, unless label-based filtering is configured in notification policies. In Zabbix channels it does not show up as a problem: it updates the d8alerts.sender.heartbeat item instead (see Zabbix channel configuration).

To avoid cluttering the alert list, DeadMansSwitch is hidden from the output of the d8 k get clusterobservabilityalerts (list/watch) command. To retrieve it directly, use the following command:

d8 k get clusterobservabilityalert deadmansswitch

Disabling this alert is not recommended, but if necessary it can be disabled manually using the deadMansSwitch.enabled parameter.

If disabled manually, the PrometheusUnavailable alert is not created.

PrometheusUnavailable

PrometheusUnavailable (formerly MissingDeadMansSwitch) is an alert that fires if DeadMansSwitch is missing for more than 2 minutes.

This indicates a problem in the alert delivery pipeline. Possible reasons include:

  • Prometheus is unavailable.
  • Communication between Prometheus and Alertmanager is broken.
  • Another issue prevents alert delivery.

The PrometheusUnavailable alert is a system alert and is displayed both in the Deckhouse web UI and in the output of the d8 k get clusterobservabilityalerts command.