The module lifecycle stage: Experimental
The module has requirements for installation
This page is the canon for release notes of the ai-models module. Every
release block lists only what matters to a cluster administrator or a namespace
user; the changelog shipped with the release image is a short projection of the
same content.
v0.1.0
Module: ai-models · Version: v0.1.0 · Date: 2026-08-06
Highlights
- Model catalog can now be filled from an external provider: a
ModelCatalogSourceindexes a Hugging Face catalog into the cluster, and the indexed entries are searchable through the catalog API without registering every model by hand. - Catalog entries carry provider metadata and inference sizing facts, so a model can be evaluated before anything is downloaded into the cluster.
- The
ai-inferencemodule resolves a model reference through an internal lookup API instead of reading catalog custom resources. - Catalog, distribution and upload endpoints are routable through Gateway API, next to the existing Ingress routing.
- The module survives a slow or contended cluster start instead of crashlooping, and its
NodeCacheruntime is admissible under the restricted Pod Security Standards profile. - The module says up front whether it can deliver models in this cluster at all: the storage prerequisites of the configured delivery mode are evaluated before any model or workload exists, and a volume the cluster cannot provide ends with a stated cause instead of an unbounded wait.
- Model storage is a fact the module reports rather than a number a consumer computes: occupancy and budget are served per scope, and every catalog entry carries a pre-download verdict on whether a download fits.
New features
ModelCatalogSourcefor the Hugging Face provider: catalog indexing runs on source events and on a timer, large catalogs are indexed incrementally with a resume cursor, and provider sync starts only after an explicit annotation on the source.- Catalog browse results expose descriptive provider metadata and inference sizing facts derived from the model configuration.
GET /api/internal/v1/models/lookupreturns model facts for in-cluster consumers by model reference.HTTPRouteresources for the catalog and distribution endpoints, and a separate route for the upload endpoint that streams without a request body size cap.- A module-level storage-prerequisites check states whether the module can deliver models in this cluster, independently of whether any
Model,ClusterModelor workload exists: the verdict and its reason are exported as thed8_ai_models_storage_prerequisites_satisfiedmetric with theD8AIModelsStoragePrerequisitesNotSatisfiedalert over it, and logged by the controller at startup for clusters without monitoring. It is re-evaluated on every scrape, so a fixed cluster clears the signal without a restart, and the module stays enabled while degraded — catalog browse, sources and the read APIs keep working. GET /api/catalog-import/v1/storage-summaryreports the occupancy of the module-owned model store, split into namespaced and cluster copies, together with the storage budget where the module has one. What the module cannot determine is absent from the response instead of being reported as a zero:capacityKnownandusageKnownsay which half is a fact.- Catalog entries carry a pre-download verdict:
downloadVerdictisFeasible,NotFeasibleorEstimateUnavailable, withrequiredByteswhen known, computed with the same reservation mechanic a real download applies. An entry whose size is not known is never blocked on an estimate the module does not have. - Storage accounting reports its own state: gauges for the inventory pass health and its last success, and for the accounting ledger’s presence and confirmation, with alerts over them — a download refused because occupancy is unknown is no longer visible only as one log line per minute.
- The whole public read surface of the module is granted at the
Useraccess level:Model,ClusterModel,ModelCatalogSourceand the catalog import projection used to browse a provider catalog. Previously a user could read models but gotForbiddenon the catalog sources they have to name inspec.source.catalog.sourceName, and browsing requiredClusterAdmin.Editor,ClusterEditorandClusterAdminare write-only deltas on top, because Deckhouse access levels are cumulative.ClusterModelandModelCatalogSourceare cluster-scoped, so a level granted withlimitNamespacesstill cannot read them — grant it cluster-wide when a user needs the shared catalog.
Improvements
- The catalog API is generated from an OpenAPI 3 contract, published as an OpenAPI 3 document: closed vocabularies became typed enums, timestamps became typed date-time fields, and the gateway failure responses 502 and 503 are part of the contract.
- Model catalog indexing keeps provider entries per source, so several sources no longer overwrite each other’s index entries.
- The storage summary is served from a five-second snapshot with a single-flight refresh, so listing the catalog no longer carries a live read of the whole accounting ledger on every browse request.
- Reading the provisioner’s explanation for an unbound volume is scoped to the claim by a server-side field selector and bounded to one page, instead of listing a whole namespace uncached on every requeue.
- Every documented condition reason is one the module actually emits: the reasons that were declared but never reported are removed, and the docs state which surface each vocabulary lives on — the
Modelconditions or the workload’s delivery report.
Fixes
- The catalog read API keeps serving the stored catalog while a source is not ready, instead of failing browse requests with an error.
- Shared volume delivery no longer overloads a cluster with parallel materialization jobs:
maxConcurrentMaterializationscaps how many run at once and defaults to 2. - The internal model API serves a certificate issued by the platform CA for its in-cluster DNS name, so callers no longer have to skip TLS verification.
- Module tags are valid again, so the module installs on a freshly built release.
- The controller records Kubernetes events for model processing and republishes a corrupt artifact instead of leaving the model in a stuck state.
- Module components no longer crashloop while starting under load: startup is gated on
/healthzthrough a startup probe, the catalog schema migration is off the readiness path, andGOMEMLIMITstays below the container memory limit. - A stopped delivery no longer destroys a materialized volume. A delivery that could not proceed — a referenced model that briefly lost
Ready, a workload blocked by its own contract — released every shared volume of its owner, after which the unreferenced-set sweep collected a claim that was already bound and carried the downloaded model. Stopping a delivery now keeps the claims that are already bound, together with the references that hold them. - An unprovisionable shared volume ends with a stated cause instead of an indefinite
SharedPVCClaimPending: the provisioner’s own failure is surfaced verbatim asSharedPVCProvisioningFailed, a claim still unbound past a bounded wait becomesSharedPVCProvisioningTimedOut, a class that defers binding is named at once instead of after the deadline, and a volume that lost itsPersistentVolumehas its own reason. No report is latched, so a cluster fix returns delivery to ready. - Storage accounting never accounts against lost state: an absent or emptied ledger reads as “occupancy unknown” rather than “nothing is stored”, new downloads are refused until an inventory pass confirms the ledger against the artifacts actually stored, and a publish commit that re-creates the ledger to record one artifact no longer passes it off as confirmed.
- An inventory pass attempts every artifact and joins the failures, so one artifact whose accounting cannot be written no longer hides the rest of the inventory while blocking confirmation for all of it.
- The storage verdict is reported by every replica as soon as it starts, instead of waiting for the leader lease — the report only reads and logs, and its entire value is being early.
Security updates
- Findings from the CVE scan are cleared across all Go modules of the module, including the update of
golang.org/x/netto 0.56.0. - The
NodeCacheCSI runtime is admissible under the restricted Pod Security Standards profile: containers drop all capabilities, the node/devtree is no longer mounted, and the single privilege the CSI node plugin still needs is declared through aSecurityPolicyException. - The report of why a shared volume is unbound only accepts an Event whose
involvedObjectUID matches the claim exactly. Creating Events is part of ordinary edit access in a workload namespace and the claim name is derived deterministically, so without the UID match a namespace user could put text of their choosing into the module’s own report. The Events access this report needs is read-only and granted only forSharedPVCdelivery.
Breaking changes
- The
DownloadTaskandModelDeliverycustom resources are removed, together with theModelandClusterModelfields that were declared and never filled:spec.availability,spec.metadata,status.state,status.discoveredandstatus.downloadTasks. No controller ever created, read or wrote any of them, so nothing a cluster does today changes — but an object or a manifest that mentions them is no longer part of the module’s API.status.phaseandstatus.resolvedare the lifecycle and enrichment fields that work, and they are no longer marked deprecated in favour of the removed ones. The internal model lookup contract drops its always-emptystatefield for the same reason.
Upgrade notes
- The
downloadtasks.ai.deckhouse.ioandmodeldeliveries.ai.deckhouse.iodefinitions stay registered in a cluster that installed an earlier version: the platform applies the definitions a module ships but does not delete the ones it stops shipping. They keep appearing inkubectl api-resourceswith whatever objects were created under them, and the controller no longer has any access to them. Nothing breaks if they are left in place; remove them when convenient withd8 k delete crd downloadtasks.ai.deckhouse.io modeldeliveries.ai.deckhouse.io, which also deletes any objects of those types. - A
StorageClasswithvolumeBindingMode: WaitForFirstConsumercannot serveSharedPVCdelivery: delivery waits for the volume to bind before it creates the volume’s first consumer, so such a volume never binds. This used to surface as an eternalSharedPVCClaimPendingand is now reported up front as an unsatisfied prerequisite. Select a ReadWriteMany-capable class withvolumeBindingMode: ImmediatethroughaiModels.delivery.sharedPVCStorageClassName, the Deckhouse global storage class settings, or a single Kubernetes defaultStorageClass.
Documentation
- The Hugging Face provider catalog is described in the user and administration guides, including request and response examples for the catalog browse API.
- The administration guide states the storage prerequisites up front: what each delivery mode requires, what each verdict reason means, and how to read the storage summary and the per-entry download verdict.
- The administration guide lists what each access level reads and writes, states that the levels are cumulative, and says that the two cluster-scoped kinds need a cluster-wide binding — including for the browse API, which answers 403 to a namespace-limited subject at the right access level.
- Operational facts that used to be answered in chat are on the pages now: a
Deploymentgets a secondReplicaSeton delivery and the superseded one is cleaned up by the controller, the storage preflight coversvolumeBindingModewhileReadWriteManyis confirmed by the first volume ordered, one bucket serves one cluster, andspec.source.catalog.nameis theClusterModelname in the publishing cluster.
Dependencies
deckhouse_lib_helm: 1.72.0 → 1.72.10.- The module requires Deckhouse 1.75.0 or newer and Kubernetes 1.34 or newer.