The module lifecycle stage: Generally available version
The module has requirements for installation
Limitations of the built-in Service load balancer
In Kubernetes, the Service resource handles the internal and external load balancing of requests. It routes requests between the application’s worker pods and excludes failed instances from load balancing. The readiness probes defined in the specs of the containers running in a pod make sure that the pod can handle incoming traffic.
The built-in Service load balancer is suitable for most cloud application tasks, but it has two limitations:
- If at least one container in a pod fails the readiness probe, the entire pod is marked as
NotReadyand is excluded from the load balancing of every service it belongs to. - You can only define one probe for each container, so it is not possible to create independent probes to check, for example, whether reads and writes are available.
The following are examples of scenarios where the capabilities of the regular load balancer are inadequate:
- Database:
- Runs as a service made up of three pods,
db-0,db-1, anddb-2. Each pod contains one container with a running database process. - You would like to create two Services,
db-writefor writing anddb-readfor reading. - Read queries must be load balanced across all pods.
- Write queries are only routed to the pod the database itself designates as master.
- Runs as a service made up of three pods,
- Virtual machine:
- The pod contains a single container running the
qemuprocess, which acts as a hypervisor for the guest virtual machine. - The guest virtual machine is running some independent processes, such as a web server and an SMTP server.
- You would like to create two Services,
webandsmtp, and define a separate readiness probe for each service.
- The pod contains a single container running the
ServiceWithHealthchecks load balancer capabilities
Unlike the regular load balancer, in which readiness probes are tied to the container state, ServiceWithHealthchecks allows you to set up active probes on individual TCP ports. In this way, each of the load balancers dealing with the same pod can operate independently of the others: a pod that fails the probes of one balancer is still published by the others.
Probes are grouped by the port they check, not by the resource. When a probe fails, the pod stops receiving traffic on that port only, and the other ports of the same resource keep working. When several probes check the same port, all of them have to pass for that port to receive traffic.
The targetPort of a probe decides which port it checks:
- If
targetPortmatches thetargetPortof one ofspec.ports, the probe checks that port only. - If
targetPortmatches none of the ports the resource exposes — a dedicated health port, for example — the probe applies to the pod as a whole, and its failure stops traffic on every port of the resource. - If no probe checks a port, that port publishes endpoints based on pod readiness alone, just like a plain Service.
You can configure this balancing method using the ServiceWithHealthchecks resource:
- Its specification is the same as that of a regular
Service, with an additionalhealthchecksection that contains a set of probes. - Currently, three types of probes are supported:
TCP— a regular probe that establishes a TCP connection.HTTP— a probe that sends an HTTP request and waits for a specific response code.PostgreSQL— a probe that sends an SQL query and waits for it to complete successfully.
In Deckhouse CE, BE, SE and SE+, only TCP probes are available. Creating or changing HTTP and PostgreSQL probes is forbidden, and such probes remaining from another edition are ignored by the agent.
Examples can be found in the documentation.
How the ServiceWithHealthchecks load balancer works
The load balancer is made up of two components:
- the controller, which runs on the cluster master nodes and manages
ServiceWithHealthchecksresources; - the agents, which run on every cluster node and probe the pods scheduled to their own node.
The ServiceWithHealthchecks load balancer is designed to be CNI implementation agnostic. It uses the built-in Kubernetes Service and EndpointSlice resources:
- When a
ServiceWithHealthchecksresource is created, the controller creates a Service of the same name in the same namespace with an emptyselectorfield. The empty selector keeps the built-in endpoint controller from creatingEndpointSliceobjects for that Service, leaving them to the agents. - The agent on each node probes the pods of the resource that are scheduled to its node and publishes the ones that pass as
EndpointSliceobjects bound to that Service. A slice lists the addresses of those pods together with the port traffic is sent to, which is thetargetPortofspec.ports— not the port a probe connects to. The two are configured separately and do not have to be the same. - The agent publishes one
EndpointSliceper port of the resource per node, named<resource>-<port>-<node>, each containing the pods that passed the probes for that port. A resource that exposes a single unnamed port publishes one slice per node, named<resource>-<node>. A node name too long for the resulting object name is replaced by a digest of itself. - kube-proxy, or the CNI, resolves the Service through all of its slices and balances traffic across the published addresses on every node of the cluster.
Migrating from a Service to a ServiceWithHealthchecks resource, for example within the framework of CI/CD, should not cause difficulties. The ServiceWithHealthchecks specification basically repeats the Service specification, but contains an additional healthcheck section. During the lifecycle of the ServiceWithHealthchecks resource, a service of the same name is created in the same namespace in order to direct traffic to workloads in the cluster in the usual way (kube-proxy or CNI).