The module lifecycle stage: Preview
The module has requirements for installation
v0.9.10
Release date: 2026-09-24
A volume that loses a diskful copy now heals by itself under a resync limit that is on by default, a single volume can be paused without stopping the controller, and the Go binaries are rebuilt against patched grpc and x/crypto.
Highlights
Changes in this release:
- Layout convergence re-creates a lost diskful copy of a
ReplicatedVolumeby itself instead of stopping atTransitionUnsupportedand waiting for a replica to be created by hand. resyncSlotsPerNodeis on by default (4), so a node going away does not start every resync of the volumes it held at once.spec.maintenance: NoResourceReconciliationon aReplicatedVolumepauses reconciliation of that volume, its replicas and its DRBD resources while every other volume keeps reconciling.
New features
This release adds:
ReplicatedVolumehasspec.maintenance. WithNoResourceReconciliation, the controller takes no action on the volume: no membership or layout decision, no replica or attachment change, no configuration derivation and no deletion handling. The pause cascades to the volume’s replicas and their DRBD resources; a DRBD resource paused by hand is left as it is.Ready,ResilientandAttachedReadykeep being published, the rest of the status freezes, and the pause is reported by theUnderMaintenancecondition and theMaintenancecolumn ofkubectl get replicatedvolume -o wide. A paused volume does not finish deleting: it and its PVC stayTerminatinguntil the field is cleared.- A
ReplicatedVolumewhose layout has fewer diskful copies than intended gets the missing copy created and resynced automatically, for example r1 → r2, r1 → r3, or a copy lost with its node. Changes that need a replica removed are still not made automatically: shrinking the copy count, replacing a copy whose member is alive, and r2 → r3 while a surplus tie-breaker is present — delete the tie-breaker, and the copy is then healed. When the copy cannot be placed, the volume reportsCannotConvergewith the scheduler’s reason.
Improvements
This release improves:
- The
resyncSlotsPerNodedefault changes from 0 (no limit) to 4, so the healing load after a node failure is bounded. A value set explicitly in the module config is kept; 0 still switches the limit off. - The controller’s health probe and loopback metrics move from ports 4271 and 4272 to 8081 and 127.0.0.1:8082, and the CSI controller’s livenessprobe no longer uses port 4261. These ports belong to the 4200–4299 range reserved for modules that listen on the host network, and neither Pod runs there. The Pods are restarted by the update.
Fixes
This release fixes:
- With the resync limiter on, several volumes could choose the same node to receive or send a resync at once, because a placement stayed invisible to other workers until it came back through the watch; the limit per node was exceeded at the start of every wave. Both ends of a transfer are now reserved before the placement is written and released as soon as the replica stops owing data.
- On a new cluster, the node label
storage.deckhouse.io/sds-replicated-volume-nodewas set only on nodes already eligible for aReplicatedStoragePool. That label is what brings up thesds-node-configuratoragent, which creates theBlockDeviceresources anLVMVolumeGroupis built from, so no pool could get eligible nodes in the first place. Nodes matchingdataNodes.nodeSelectornow get the label directly.
Security updates
Security updates in this release:
google.golang.org/grpcis raised from v1.82.1 to v1.83.2 in the CSI driver,linstor-csi,linstor-affinity-controllerandlinstor-scheduler-extender, closing CVE-2026-84303, CVE-2026-84304 and CVE-2026-84445.golang.org/x/cryptois raised from v0.53.0 to v0.57.0 in the module’s hooks (go-hooks) andlinstor-scheduler-extender, closing CVE-2026-56854 (authentication bypass by source address inx/crypto/ssh), CVE-2026-56855 and CVE-2026-78662.
Breaking changes
Changes that affect backward compatibility:
evict.shandreplicas_manager.share installed in/var/lib/deckhouse/sbininstead of/opt/deckhouse/sbin, and the old copies are removed from nodes. Runbooks and scripts that call them by the old path have to use the new one.
Upgrade notes
Before upgrading, note the following:
- Nothing has to be done by hand. After the update, the node configuration moves
evict.shandreplicas_manager.shto/var/lib/deckhouse/sbin; update any runbook that calls/opt/deckhouse/sbin/evict.sh. - If
resyncSlotsPerNodewas not set, the limit of 4 applies after the update, and volumes that were short of a diskful copy start healing within that limit. To keep resyncs unlimited, setresyncSlotsPerNode: 0.
Docs
Documentation changes:
- The usage guide describes recovering a lost replica, which no longer needs a manual step, and the
D8ReplicatedVolumeLayoutDegradedalert no longer walks through creating a replica by hand.
v0.9.9
Release date: 2026-09-21
A ReplicatedVolume now reports where its replicas sit, -o wide gained the two columns the configuration and layout pictures were missing, and switching the resync limiter on no longer drives the controller into its memory ceiling.
Highlights
Changes in this release:
- A
ReplicatedVolumereports its placement: theScheduledandSatisfyEligibleNodesconditions are set for the first time, andstatus.eligibleNodesViolationsnames the replicas placed on nodes their storage pool does not allow, with the reason for each. kubectl get replicatedvolume -o wideshowsCONVERGEDandCONFIGURATIONREADY, so a volume whose layout is stuck or whose configuration is invalid is visible in the list instead of only inside the object.- Switching
resyncSlotsPerNodeon no longer costs the controller the whole cluster in memory: a cluster of thousands of volumes used to push it into its pod limit.
New features
This release adds:
Scheduledreports whether every replica of a volume has landed on a node, aggregated over the non-deleting replicas the scheduler places —Accessreplicas, which follow their consumer, are excluded. It is orthogonal to availability and does not enterReady: a volume with a misplaced replica keeps serving I/O.SatisfyEligibleNodesreports whether the placement agrees with the storage pool’s eligible nodes, andstatus.eligibleNodesViolationslists the offending replicas with their node and a reason. Beyond the per-replica verdict it checks what only the member set can show: underTransZonala node carrying no zone, a zone the pool has no eligible nodes in, and all voters collapsed into a single zone; underZonalevery voter outside the home zone.- A
ReplicatedStorageClasscounts those volumes instatus.volumes.inConflictWithEligibleNodesand sets itsVolumesSatisfyEligibleNodescondition accordingly — the counter was structurally pinned at zero while the volume condition was never written.
Improvements
This release improves:
- Two wide print columns are back in the
ReplicatedVolumelist:CONVERGED(theMembershipLayoutConvergedcondition, the only list-level view of a membership that will not move without a human) andCONFIGURATIONREADY(whether the configuration itself is valid and current against the storage class, the halfCONFIGUREDdoes not answer). Both are wide-only, so the default output is unchanged.
Fixes
This release fixes:
- The resync limiter no longer reads the whole cluster for every decision. Each budget decision used to materialise every
ReplicatedVolumeand everyReplicatedVolumeReplica, and ten workers doing that at once exceeded the controller pod’s memory allowance; the listings are now narrowed by field indexes and the queue is built one claim at a time, so a candidate costs a few hundred bytes rather than its replicas. The decisions themselves are unchanged. - Metrics of the CSI controller and the CSI node driver are scraped again. Their
kube-rbac-proxysidecars authorise every scrape through aTokenReviewand aSubjectAccessReview, and thecsiServiceAccount was bound to neither, so the proxy answered 401 while the pods stayedReady— the only symptom wasup == 0on both jobs and aTargetDownthat never cleared. - The same binding was missing for the legacy control plane’s controller: its metrics proxy is rendered, the
ServiceMonitorselects the Service, and every scrape was refused by the API server.
Upgrade notes
Before upgrading, note the following:
- Nothing has to be done by hand, and no volume is moved: the two placement conditions only report what is already there.
- A storage class whose volumes have replicas on non-eligible nodes can leave
Readyfor the first time after the update — forPartiallyAlignedunder theManualresolution strategy, forRollingOutwhere conflict resolution is enabled. The placement was already wrong; it is now reported. Which replicas are meant is instatus.eligibleNodesViolationsof the volumes the class counts.
Docs
Documentation changes:
- The Replication r3 → r2 migration guide names the print columns that exist: progress per volume is watched through
LAYOUTand, in-o wide,CONVERGED.
v0.9.8
Release date: 2026-09-14
The module’s certificates are now renewed and rolled out on their own, the false expiry alert is gone, and the controller publishes metrics behind a kube-rbac-proxy.
Highlights
Changes in this release:
- Module certificates are re-signed 45 days before they expire and the workloads using them are restarted automatically. Until now nothing renewed a certificate at all: it was regenerated only when its Secret went missing or empty.
D8LinstorCertificateExpiringIn30dno longer fires for certificates that are not the module’s own. The hook labelled every Secret in the module namespace, including the Let’s Encrypt ones copied there bysecret-copier.- The controller serves metrics behind a
kube-rbac-proxy, with aServiceMonitorfor Prometheus.
New features
This release adds:
- The controller publishes metrics: it binds them on
127.0.0.1and akube-rbac-proxyin the same pod serves them over TLS, so reaching them takesgetondeployments/prometheus-metricsand the grant goes to the Prometheus scraper alone. AServiceMonitoris created where the kind exists, and the ports are settable through theoldControllerMetricsPortandoldControllerMetricsProxyPortmodule values.
Improvements
This release improves:
- The controller logs through the shared logger in the structured format the rest of the fleet uses, instead of its own copy of an older logging package.
Fixes
This release fixes:
- Certificates are renewed without manual intervention. Every certificate group now keeps its CA, certificate and private key in its own Secret, and a certificate is re-signed by the same CA 45 days before it expires; the CA lives ten years and is replaced only once it can no longer cover a full certificate lifespan. Renewal takes effect on a restart, and a new hook restarts the workloads that reference a renewed certificate Secret.
- The expiring-certificate hook now looks only at the module’s own Secrets, strips its label from the rest and no longer panics on a Secret without labels. Before this, every Secret in the module namespace was labelled, so certificates copied in by other modules tripped
D8LinstorCertificateExpiringIn30devery few weeks. - The
linstor-scheduler-admission-certsSecret, left over from a component removed earlier, no longer forces the webhook certificate group to be re-issued on every hook run — which re-minted the CA and rewrote the webhookcaBundleeach time. It is dropped from the group and from the openapi schema, and a hook deletes the leftover Secret. - A volume whose filesystem the CSI driver creates itself after a failed mount is formatted with
mkfs.ext4 -m0, so it no longer hands 5% of its space to the superuser.
Upgrade notes
Before upgrading, note the following:
- Nothing has to be done by hand, and the update rotates nothing: certificate groups issued by earlier versions, whose CA key is unknown, are left untouched until their certificates go outdated.
- Workloads that need a renewed certificate are restarted by a hook on the convergence that follows the renewal. Renewing 45 days ahead leaves the previous certificate valid in the meantime.
v0.9.7
- Bugfix: the controller is granted patch on events instead of list - a repeated event write is no longer denied by RBAC
- Bugfix: the VPA is switched from the deprecated Auto update mode to InPlaceOrRecreate
- Optional resources are enabled based on the presence of the CRD in the API, not on the list of enabled modules - the checks are routed through helm_lib_api_version_exists
- Base images updated to v2.1.2 and Go to 1.26.6
- Internal build changes: CI updated, crds, docs and openapi are shipped in the release image
v0.9.6
- Correction: the CSI module service account is no longer blocked by the validating webhook when creating an RWX PersistentVolumeClaim
- Fixed links in documentation
- Updating base images and lib-helm to 1.72.13
v0.9.5
- Dependency on snapshot-controller module made optional: VolumeSnapshotClass is created only if the snapshot.storage.k8s.io group CRD is present, the rest of the functionality works without it
- Correction: only the credentials of the active image source remain in the deckhouse-registry secret
- Correction: LINSTOR no longer requests a list of all cluster CRDs when determining the database version
- Correction: LINSTOR ignores unknown object properties when working with Kubernetes API
- Correction: installing DRBD on nodes reinstalls the package if the sources are outdated and more strictly checks the version of the loaded module
- Re-issuance of certificates has been switched to a manual procedure described in the FAQ; manualcertrenewal-trigger trigger hook removed
- Updating base images, container-base and lib-helm to 1.72.12
- Fixed vulnerabilities found when scanning module images
- Internal changes in module assembly
v0.9.4
- When forwarding labels from ReplicatedStorageClass to StorageClass, labels with specified ignored prefixes are now excluded
- Updating base images and lib-helm to 1.72.0
v0.9.3
- Labels from ReplicatedStorageClass are now forwarded to the Kubernetes managed StorageClass
- Update base images, Go 1.25.10 and lib-helm 1.71.12
v0.9.2
- Added x-config-version marker to the OpenAPI schema config-values for versioning module parameters
v0.9.1
- Documentation - clarifications in the FAQ about SSDs
- Internal changes in module structure and assembly
v0.9.0
- Update DRBD version to 9.2.16. Important: nodes will need to be rebooted
- Internal changes in the module
v0.8.19
- Update base images, Go 1.25.10 and lib-helm 1.71.11
- Internal changes to the module assembly
v0.8.18
- Corrections to the module structure
v0.8.17
- CI changes: DistroPackagesProxy and env proxy in werf, improvements to CVE scans (role_name, checkout)
- Added user-authz cluster roles in templates
v0.8.16
- Added missing mount points in webhooks distroless image
- Update base images, Go and lib-helm (CVE fix)
- Clarifications in the documentation
v0.8.15
- Added a hook for clearing resources when the module is turned off
- Removed the enableThinProvisioning setting, we assume that thin provisioning is always enabled
- Module-sdk update for CVE fix, golang update
v0.8.14
- Corrections in documentation
- Reworking manifestos
- Added a ban on ordering RWX volumes for all except the virtualization module
v0.8.13
- Switch from yum to dnf for package installation on CentOS-like distributions
- Fix CVE
- Updated documentation
- Increased liveness probe timeout for linstor-controller to 60 seconds
v0.8.12
- Documentation fixes
- Fix HA-mode, now it works correctly
- Removed scheduler functionality, remaining for compatibility with old Deckhouse versions
v0.8.11
- Updated base images versions
v0.8.10
- Updated Go version to 1.24.9
- Updated lib-helm to deckhouse_lib_helm-1.64.1
v0.8.9
- Fixed satellite bind address to avoid potential external access to the cluster
v0.8.8
- Update base images to v0.5.34
- Add pull secrets to CSI controller and node
v0.8.7
- Fixed error with not created FS when creating volumes
- Documentation fixes
v0.8.6
- Added release notes
- Hooks switched from python to golang
- Docs improved
v0.8.5
- Added additional mountings for containerd v2 support
v0.8.4
- Added information about the need for snapshot-controller for module operation
v0.8.3
- Documentation fixes
- Added dependency on snapshot-controller
v0.8.2
- Certificate update hook fixes
- Removal of obsolete migration hooks
v0.8.1
- Documentation fixes (added instruction for expanding ReplicatedStoragePool to a new cluster node)
v0.8.0
- Module refactoring
- Documentation fixes
- Fixes for volume snapshot support
v0.7.4
- If topology allows, controller removes annotation for StorageClass that prohibits ordering RWX volumes
v0.7.3
- Module refactoring
- Fixed podAntiAffinity for sds-replicated-volume-controller
v0.7.2
- Changes in hooks for correct manual certificate update process
- Fixed D8NodeHighUnknownMemoryUsage alert grouping
v0.7.1
- Added CSI patch for full support of topologies specified in ReplicatedStorageClass (could be ignored)
- Added D8NodeHighUnknownMemoryUsage alert for detecting DRBD memory leak cases (report issues to team storage)
v0.6.0
- Updated DRBD to version v9.2.12, solving a number of problems (particularly improving DRBD diskless replica stability)
v0.5.1
- Fixed alert for incorrect number of resource replicas
- Fixed schedule job for Linstor database backup
v0.5.0
- Multiple minor fixes in templates, monitoring alerts and documentation
- Transition from linstor scheduler-extender to internal Deckhouse mechanisms (KubeSchedulerWebhookConfiguration)
- Migration of images to distroless
- Fixed and enhanced script for outputting drbd resources from node
v0.4.3
- Technical release. Fixes and additions to evict.sh script for resource eviction from node, fixes in templates and documentation
v0.4.1
- In evict.sh script for cleaning node from DRBD resources, AutoplaceTarget parameter is now taken into account, moved replicas will not be moved to nodes with AutoplaceTarget value equal to false
v0.4.0
- Updated golang API libraries for sds-node-configurator v0.4.0 support
- Multiple fixes in controllers and documentation
v0.3.5
- Multiple fixes and improvements in evict.sh and replicas_managers.sh (also, they are now automatically installed in /opt/deckhouse/sbin)
- DRBD now correctly builds on ALT Linux and with Linux kernel 6.5+
- Added anti-affinity rules for controller pods
- Multiple fixes in dashboard and alerts
- isDefault parameter removed; use standard k8s annotation instead
- Added liveness and readiness checks for controllers
- Backup switched to dedicated CR instead of using secrets in module namespace
- Prohibited creation of pools on ephemeral nodes
- Multiple documentation fixes
- CSI endpoint migrates from linstor.csi.linbit.com to replicated.csi.deckhouse.io
v0.2.9
- Add DRBD ports range settings
- Fix path in liveness-satellite
- Actual typo lvmVolumeGroups and thinPoolName in examples
- Add a check for a Linstor node’s AutoplaceTarget property
- Changed lvmvolumegroups to lvmVolumeGroups in russian docs
- Fix linstor satellite VPA
v0.2.8
- Add check if /etc/modules file exists
- Add liveness probe for linstor-node
- Add age field
- Prevent Node Reboots on Volume Deletion
v0.1.11
- Fix enabled script, module will not be disabled if sds-node-configurator module disappears from cluster