The module lifecycle stage: Generally available version
The module has requirements for installation
Triggers
Triggers (alerting rules) define the conditions for creating alerts when metric values deviate from expected thresholds.
Triggers are defined in rule groups as items in the spec.rules array.
If a rule contains the alert field, it is treated as a trigger and is used to create alerts.
Types of rule groups with triggers
Three types of rule groups can define triggers:
| Rule group type | Scope | Who has access |
|---|---|---|
| System rule groups (ClusterObservabilityMetricsRulesGroup) | Cluster level | DKP administrators |
| Project rule groups (ObservabilityMetricsRulesGroup) | Project (namespace) level | Users of the corresponding project |
| Standard propagated rule groups (ClusterObservabilityPropagatedMetricsRulesGroup) | Created at the cluster level and automatically available in all projects | Users of all projects |
Rule group types:
-
System rule groups (ClusterObservabilityMetricsRulesGroup): Used to define triggers for platform-level and cluster component alerts. Created and managed by DKP administrators.
-
Project rule groups (ObservabilityMetricsRulesGroup): Used to define triggers related to a specific project (namespace). Project users can create and edit them within the configured access permissions.
-
Standard propagated rule groups (ClusterObservabilityPropagatedMetricsRulesGroup): Created at the cluster level and automatically become available in all projects.
Additional labels for alerts shipped with DKP
The observability module lets you apply additional labels to alerting rules shipped with DKP.
To do this, use the ClusterObservabilityAlertAdditionalLabels resource.
Additional labels are applied only to alerts created using the ClusterObservabilityMetricsRulesGroup and ClusterObservabilityPropagatedMetricsRulesGroup resources with the heritage: deckhouse label.
To add labels to custom alerts created with ObservabilityMetricsRulesGroup or ClusterObservabilityMetricsRulesGroup, use the spec.rules.labels field.
Configuration examples:
-
Adding a label to all cluster alerts:
apiVersion: observability.deckhouse.io/v1alpha1 kind: ClusterObservabilityAlertAdditionalLabels metadata: name: all-alerts spec: alertSelector: matchExpressions: - key: alertname operator: Exists additionalLabels: example-label-name: example-label-value -
Adding the
severity=Infolabel to alerts withseverity_levelof 7, 8, and 9:apiVersion: observability.deckhouse.io/v1alpha1 kind: ClusterObservabilityAlertAdditionalLabels metadata: name: severity-info-low-priority spec: alertSelector: matchExpressions: - key: severity_level operator: In values: ["7", "8", "9"] additionalLabels: severity: Info -
Adding the
team=customlabel to alerts with the specified names:apiVersion: observability.deckhouse.io/v1alpha1 kind: ClusterObservabilityAlertAdditionalLabels metadata: name: deckhouse-team-routing spec: alertSelector: matchExpressions: - key: alertname operator: In values: - D8CNIMisconfigured - D8DeckhouseIsNotOnReleaseChannel - D8DeckhouseIsNotOnReleaseChannel additionalLabels: team: custom
Trigger groups
Trigger groups are used to logically organize monitoring rules and manage their parameters at the rule set level.
Groups are convenient for combining triggers related to the same component, service, or project, as well as for applying a shared evaluation interval to all rules in the group.
Notifications
The observability module provides mechanisms for configuring alert notification delivery
and controlling access to notification channels at both the cluster and project levels.
The following delivery channels are supported:
EmailTelegramSlackWebhookExpressMessengerZabbix
Connection parameters depend on the channel type and are configured through the corresponding Kubernetes resource.
Protection of channel credentials
Credential fields are protected by the x-kubernetes-sensitive-data marker: without rights to the
<resource>/sensitive subresource, the API server returns them as <omitted> in get/list/
watch and masks them in the audit log. /sensitive access is granted to anyone who can create/edit the channel itself; levels without write access do not get it.
Requires the CRDSensitiveData feature gate on kube-apiserver — enabled by default starting with DKP 1.77; on earlier versions the marker has no effect and credentials stay readable.
Types of notification channels
Three types of notification channels are supported:
| Channel type | Scope | Who can create |
|---|---|---|
| System channels (ClusterObservabilityNotificationChannel) | Cluster level | DKP administrators |
| Project channels (ObservabilityNotificationChannel) | Project (namespace) level | Users of the corresponding project |
| Standard propagated channels (ClusterObservabilityPropagatedNotificationChannel) | Created at the cluster level and automatically available in all projects | DKP administrators |
Channel types:
-
System channels (ClusterObservabilityNotificationChannel): Used for cluster-level notification delivery. Available in the Deckhouse web UI under “System” → “System management” → “Monitoring” → “Notification settings” → “Notification channels”.
-
Project channels (ObservabilityNotificationChannel): Allow configuring notification delivery within a specific project. Available in the corresponding project in the Deckhouse web UI under “Monitoring” → “Notification settings” → “Notification channels”.
-
Standard propagated channels (ClusterObservabilityPropagatedNotificationChannel): Created at the cluster level and automatically become available in all projects for notification delivery. Use the ClusterObservabilityPropagatedNotificationChannel resource or the
d8CLI utility to create them.
Webhook channel HTTP client configuration
For channels with spec.type: Webhook, you can additionally configure outbound HTTP client settings in spec.webhook.httpConfig.
There are 3 mutually exclusive authentication options:
basicAuthauthorizationoauth2
Required and optional fields
- Required for a Webhook channel:
spec.webhook.url spec.webhook.httpConfigis optional- If
oauth2is used,oauth2.clientIdandoauth2.tokenUrlare required
In addition to authentication, httpConfig supports:
- transport options (
enableHttp2) - proxy settings (
proxyUrl,noProxy,proxyFromEnvironment,proxyConnectHeader) - TLS settings (
tlsConfig) - custom headers (
httpHeaders)
For OAuth2, there are two configuration levels:
httpConfig.proxy*andhttpConfig.tlsConfigapply to webhook delivery requestshttpConfig.oauth2.proxy*andhttpConfig.oauth2.tlsConfigapply to OAuth2 token requests
The following fields are deprecated and ignored for security reasons. They are still accepted by the CRD schema for backward compatibility, but they have no effect:
followRedirects— redirects are never followed, because a redirect makes Alertmanager fetch a location chosen by the notification target.- the file-based fields
passwordFile,credentialsFile,clientSecretFile,tlsConfig.caFile,tlsConfig.certFile,tlsConfig.keyFileandhttpHeaders.<name>.files— the files are no longer read from the Alertmanager container. Use the corresponding inline field (password,credentials,clientSecret) orhttpHeaders.<name>.values/.secretsinstead.
Example with basicAuth:
apiVersion: observability.deckhouse.io/v1alpha1
kind: ClusterObservabilityNotificationChannel
metadata:
name: webhook-channel-basic-auth
spec:
type: Webhook
webhook:
url: https://hooks.example/webhook
httpConfig:
basicAuth:
username: notify-user
password: notify-secretExample with authorization:
apiVersion: observability.deckhouse.io/v1alpha1
kind: ClusterObservabilityNotificationChannel
metadata:
name: webhook-channel-authorization
spec:
type: Webhook
webhook:
url: https://hooks.example/webhook
httpConfig:
authorization:
type: Bearer
credentials: opaque-tokenExample with oauth2:
apiVersion: observability.deckhouse.io/v1alpha1
kind: ClusterObservabilityNotificationChannel
metadata:
name: webhook-channel-oauth2
spec:
type: Webhook
webhook:
url: https://hooks.example/webhook
httpConfig:
oauth2:
clientId: my-client
clientSecret: my-secret
tokenUrl: https://idp.example/token
scopes:
- read
- write
endpointParams:
audience: myappZabbix channel configuration
For channels with spec.type: Zabbix, Alertmanager pushes alerts to a Zabbix Server or Zabbix Proxy
in push mode, using the Zabbix sender (trapper) protocol.
Required and optional fields
- Required:
spec.zabbix.server— the address of the Zabbix Server or Zabbix Proxy. - Required:
spec.zabbix.clusterName— an arbitrary string that becomes the technical name of the Zabbix Host that owns the pushed items and the imported template. Pick any value that’s unique among every cluster/channel pointed at the same Zabbix instance (for example, your cluster’s public domain, if you have one, or any other identifier your team already uses for this cluster). spec.zabbix.port— the sender protocol port (default10051).spec.zabbix.hostGroup— the name of the Zabbix host group the cluster’s Host is placed into. Defaults toDeckhouseif unset.- The sender connection to Zabbix is encrypted by default.
spec.zabbix.tlsoptionally configures it further (caFile,certFile,keyFile,serverName,insecureSkipVerify— for a private CA or a self-signed certificate; a public CA needs no configuration). Zabbix’s PSK encryption mode is not supported.spec.zabbix.disableTLS: trueopts out entirely, for a Zabbix Server/Proxy with no TLS on its trapper port. spec.zabbix.api— optional; enables automatic import of the bundled template via the Zabbix API (see “Automatic import via the Zabbix API” below). If unset, only the manual-importConfigMapis produced.
When an alert stops firing
- A resolved alert is sent once more with the value
0, which clears its trigger immediately. - The next notification no longer lists the alert in the discovery data, so Zabbix marks its item
as lost and deletes it after the discovery rule’s
lifetimeof 1 hour; the alert then disappears from Latest data too. - As a fallback, the trigger also stops firing once its item has received nothing for 5 minutes (for example when the alert was silenced and no resolved notification was sent).
Because the freshness check relies on alerts being re-sent, spec.notification.repeatInterval for a
Zabbix channel is capped at 2m (a longer value is silently reduced); the default of 30s is fine.
spec.alert.groupByLabels is ignored for Zabbix channels: all alerts matched by the policy are sent
together in every notification, which is what the Zabbix discovery rules require. Which alerts are
sent is still controlled by spec.alert.selector.
The d8alerts.sender.heartbeat item is kept fresh by the platform’s always-on DeadMansSwitch
alert, routed to the channel like any other alert — through its ObservabilityNotificationPolicy
selector. A narrow selector (e.g. severity="critical") needs an explicit
alertname="DeadMansSwitch" match added to keep the heartbeat working. Zabbix uses nodata() on
that item to detect if the path goes down.
The heartbeat item exists only in the template of a ClusterObservabilityNotificationChannel. A
namespaced ObservabilityNotificationChannel and a ClusterObservabilityPropagatedNotificationChannel
never receive DeadMansSwitch, so their templates have no heartbeat item.
Registering the template in Zabbix: manual or automatic
The bundled template (zabbix_export format) already includes the cluster’s Host and Host Group
(spec.zabbix.clusterName/hostGroup), so importing it leaves nothing else to configure by hand.
There are two ways to get it into Zabbix:
Manual import (always available)
The controller always renders the template into a ConfigMap for the channel, regardless of
spec.zabbix.api. The name and namespace depend on the channel kind:
| Channel kind | ConfigMap name | Namespace |
|---|---|---|
ObservabilityNotificationChannel |
<channel-name>-zabbix-template |
same as the channel |
ClusterObservabilityNotificationChannel |
cluster-<channel-name>-zabbix-template |
module namespace |
ClusterObservabilityPropagatedNotificationChannel |
propagated-<channel-name>-zabbix-template |
module namespace |
Extract it and import it into Zabbix under Data collection → Templates → Import, for example
for a project channel named my-zabbix-channel in namespace my-namespace:
kubectl get configmap my-zabbix-channel-zabbix-template -n my-namespace \
-o jsonpath='{.data.template\.yaml}' > d8alerts-sender.yamlAutomatic import via the Zabbix API (optional)
Setting spec.zabbix.api makes the controller import the exact same rendered template into Zabbix
directly, via its JSON-RPC API (configuration.import), in addition to (not instead of) the
ConfigMap above:
spec:
type: Zabbix
zabbix:
server: zabbix.example.com
clusterName: my-cluster
hostGroup: Deckhouse
api:
url: https://zabbix.example.com/api_jsonrpc.php
token: <zabbix-api-token>spec.zabbix.api.url— address of the Zabbix JSON-RPC API; if the scheme is omitted,https://is assumed. This is a different address thanspec.zabbix.server: a Zabbix Proxy has no API of its own, soapi.urlmust point at the central Zabbix Server. Requires Zabbix ≥6.4.spec.zabbix.api.token— a Zabbix API token, stored in this resource in plain text. Create a dedicated token with a narrowly-scoped role (limited to the relevant template group and host group) rather than reusing a Super Admin account.spec.zabbix.api.tlsConfig— optional TLS settings for the connection to the Zabbix API (mirrorsspec.zabbix.tls, but the certificate/key paths must be available inside theobservability-controllercontainer instead of thealertmanagercontainer).
If the import fails, check the observability-controller logs for failed to import zabbix template via api. Typical causes: the Zabbix API is unreachable, the token is invalid or lacks permission on
the relevant template or host group, or the Zabbix version is lower than 6.4. The import is retried
automatically on every reconcile (e.g. after editing the channel, or the next time the controller
restarts), so a transient failure clears itself without any action.
Choosing the Zabbix server address
Prefer a stable DNS name over a literal IP for spec.zabbix.server and spec.zabbix.api.url — for
in-cluster Zabbix, a Service DNS name rather than a node or pod IP. The admission webhook emits a
non-blocking warning for literal IPs.
Notification policies
Notification policies define which channel should be used to deliver notifications for an alert (or a group of alerts).
| Policy type | Description | How to configure |
|---|---|---|
| System notification policies | Used to configure delivery rules for system alerts. System policies can use only system notification channels. Available in the Deckhouse web UI under “System” → “System management” → “Monitoring” → “Notification settings” → “Notification policies”. | Use the ClusterObservabilityNotificationPolicy resource. |
| Project notification policies | Used to configure delivery rules for project alerts. Project policies can use project or standard cluster channels, but not system notification channels. Available in the corresponding project under “Monitoring” → “Notification settings” → “Notification policies”. | Use the ObservabilityNotificationPolicy resource. |
Notification silencing
In situations where notifications are expected in advance (for example, during planned maintenance or testing),
the observability module allows disabling notification delivery for alerts matching specified conditions.
| Silence type | Description | How to configure |
|---|---|---|
| System notification silences | Used to configure silencing rules for system alert delivery. Available in the Deckhouse web UI under “System” → “System management” → “Monitoring” → “Notification settings” → “Notification silencing”. | Use the ClusterObservabilityNotificationSilence resource. |
| Project notification silences | Used to configure silencing rules for project alert delivery. Available in the corresponding project under “Monitoring” → “Notification settings” → “Notification silencing”. | Use the ObservabilityNotificationSilence resource. |
Alerts
The observability module provides access control separation for cluster-level and project-level alerts
and allows viewing the list of active and resolved alerts.
Active alerts are grouped by severity level:
- critical (
critical, S1–S3) - warning (
warning, S4–S6) - informational (
info, S7–S9)
When viewing an alert, the user can see general information, labels, annotations, and a graph.
Types of alerts
Two types of alerts are supported:
| Alert type | Scope | Who has access |
|---|---|---|
| System alerts (ClusterObservabilityAlerts) | Cluster level | DKP administrators |
| Project alerts (ObservabilityAlerts) | Project (namespace) level | Users of the corresponding project |
Alert types:
-
System alerts (ClusterObservabilityAlerts): Relate to DKP cluster components. The full list of active and resolved system alerts is available in the Deckhouse web UI under “System” → “System management” → “Monitoring” → “Active alerts”.
-
Project alerts (ObservabilityAlerts): Relate to resources of a specific project (namespace). The full list of active and resolved project alerts is available in the Deckhouse web UI in the corresponding project under “Monitoring” → “Active alerts”.
DeadMansSwitch and PrometheusUnavailable alerts
DeadMansSwitch
DeadMansSwitch is a service alert that fires continuously, confirming the normal operation of Prometheus and the entire alert delivery pipeline.
If DeadMansSwitch stops arriving, the PrometheusUnavailable alert starts firing.
By default, the DeadMansSwitch alert is sent to all configured notification channels,
unless label-based filtering is configured in notification policies. In Zabbix channels it does
not show up as a problem: it updates the d8alerts.sender.heartbeat item instead (see
Zabbix channel configuration).
To avoid cluttering the alert list, DeadMansSwitch is hidden from the output of the d8 k get clusterobservabilityalerts (list/watch) command.
To retrieve it directly, use the following command:
d8 k get clusterobservabilityalert deadmansswitchDisabling this alert is not recommended, but if necessary it can be disabled manually
using the deadMansSwitch.enabled parameter.
If disabled manually, the PrometheusUnavailable alert is not created.
PrometheusUnavailable
PrometheusUnavailable (formerly MissingDeadMansSwitch) is an alert that fires if DeadMansSwitch is missing for more than 2 minutes.
This indicates a problem in the alert delivery pipeline. Possible reasons include:
- Prometheus is unavailable.
- Communication between Prometheus and Alertmanager is broken.
- Another issue prevents alert delivery.
The PrometheusUnavailable alert is a system alert and is displayed both in the Deckhouse web UI
and in the output of the d8 k get clusterobservabilityalerts command.