The module lifecycle stage: Generally available version

The module has requirements for installation

Limitations of the built-in Service load balancer

In Kubernetes, the Service resource handles the internal and external load balancing of requests. It routes requests between the application’s worker pods and excludes failed instances from load balancing. The readiness probes defined in the specs of the containers running in a pod make sure that the pod can handle incoming traffic.

The built-in Service load balancer is suitable for most cloud application tasks, but it has two limitations:

  • If at least one container in a pod fails the readiness probe, the entire pod is marked as NotReady and is excluded from the load balancing of every service it belongs to.
  • You can only define one probe for each container, so it is not possible to create independent probes to check, for example, whether reads and writes are available.

The following are examples of scenarios where the capabilities of the regular load balancer are inadequate:

  • Database:
    • Runs as a service made up of three pods, db-0, db-1, and db-2. Each pod contains one container with a running database process.
    • You would like to create two Services, db-write for writing and db-read for reading.
    • Read queries must be load balanced across all pods.
    • Write queries are only routed to the pod the database itself designates as master.
  • Virtual machine:
    • The pod contains a single container running the qemu process, which acts as a hypervisor for the guest virtual machine.
    • The guest virtual machine is running some independent processes, such as a web server and an SMTP server.
    • You would like to create two Services, web and smtp, and define a separate readiness probe for each service.

ServiceWithHealthchecks load balancer capabilities

Unlike the regular load balancer, in which readiness probes are tied to the container state, ServiceWithHealthchecks allows you to set up active probes on individual TCP ports. In this way, each of the load balancers dealing with the same pod can operate independently of the others: a pod that fails the probes of one balancer is still published by the others.

Probes are grouped by the port they check, not by the resource. When a probe fails, the pod stops receiving traffic on that port only, and the other ports of the same resource keep working. When several probes check the same port, all of them have to pass for that port to receive traffic.

The targetPort of a probe decides which port it checks:

  • If targetPort matches the targetPort of one of spec.ports, the probe checks that port only.
  • If targetPort matches none of the ports the resource exposes — a dedicated health port, for example — the probe applies to the pod as a whole, and its failure stops traffic on every port of the resource.
  • If no probe checks a port, that port publishes endpoints based on pod readiness alone, just like a plain Service.

You can configure this balancing method using the ServiceWithHealthchecks resource:

  • Its specification is the same as that of a regular Service, with an additional healthcheck section that contains a set of probes.
  • Currently, three types of probes are supported:
    • TCP — a regular probe that establishes a TCP connection.
    • HTTP — a probe that sends an HTTP request and waits for a specific response code.
    • PostgreSQL — a probe that sends an SQL query and waits for it to complete successfully.

In Deckhouse CE, BE, SE and SE+, only TCP probes are available. Creating or changing HTTP and PostgreSQL probes is forbidden, and such probes remaining from another edition are ignored by the agent.

Examples can be found in the documentation.

How the ServiceWithHealthchecks load balancer works

The load balancer is made up of two components:

  • the controller, which runs on the cluster master nodes and manages ServiceWithHealthchecks resources;
  • the agents, which run on every cluster node and probe the pods scheduled to their own node.

The ServiceWithHealthchecks load balancer is designed to be CNI implementation agnostic. It uses the built-in Kubernetes Service and EndpointSlice resources:

  • When a ServiceWithHealthchecks resource is created, the controller creates a Service of the same name in the same namespace with an empty selector field. The empty selector keeps the built-in endpoint controller from creating EndpointSlice objects for that Service, leaving them to the agents.
  • The agent on each node probes the pods of the resource that are scheduled to its node and publishes the ones that pass as EndpointSlice objects bound to that Service. A slice lists the addresses of those pods together with the port traffic is sent to, which is the targetPort of spec.ports — not the port a probe connects to. The two are configured separately and do not have to be the same.
  • The agent publishes one EndpointSlice per port of the resource per node, named <resource>-<port>-<node>, each containing the pods that passed the probes for that port. A resource that exposes a single unnamed port publishes one slice per node, named <resource>-<node>. A node name too long for the resulting object name is replaced by a digest of itself.
  • kube-proxy, or the CNI, resolves the Service through all of its slices and balances traffic across the published addresses on every node of the cluster.

Migrating from a Service to a ServiceWithHealthchecks resource, for example within the framework of CI/CD, should not cause difficulties. The ServiceWithHealthchecks specification basically repeats the Service specification, but contains an additional healthcheck section. During the lifecycle of the ServiceWithHealthchecks resource, a service of the same name is created in the same namespace in order to direct traffic to workloads in the cluster in the usual way (kube-proxy or CNI).