A load balancer is a system that receives network traffic and distributes it among multiple available targets, such as web servers, application processes, or containers. Clients connect to one service address instead of choosing a back-end machine themselves. The load balancer accepts or forwards each connection according to configured rules and the condition of the targets. This lets an application add capacity, replace servers, or survive some failures without requiring every client to learn a new destination. It is a traffic coordinator, not the source of the application’s content.
Different load balancers operate at different layers. A transport-level balancer can route TCP or UDP connections using addresses, ports, and connection state. An application-level balancer understands protocols such as HTTP and can make decisions using a hostname, URL path, header, or other request property. It might send image requests to one pool and account requests to another. Some services terminate TLS, meaning they handle the encrypted client connection and create a separate connection to the selected target. That can centralize certificate management but also changes the security boundary.
Selection algorithms answer the immediate question of where traffic should go. Round robin rotates among eligible targets. Least-connections methods favor a target handling fewer active connections, while weighted rules give stronger machines or preferred groups a larger share. Hash-based methods can choose a target consistently from information such as a client or request key. No algorithm is universally best: short uniform requests behave differently from long downloads, streaming sessions, or tasks with uneven processing cost. Operators choose and tune the method for the workload.
Health checks keep obviously failed targets out of rotation. The load balancer periodically attempts a connection or requests a defined endpoint and compares the result with success criteria. After repeated failures, it stops sending new traffic to that target; after sufficient successful checks, it can restore the target. A useful check must represent the service’s real ability to work. A superficial endpoint may report success while a critical database dependency is unavailable, whereas an overly demanding check can remove healthy capacity during a brief slowdown.
Applications also need a plan for state. If every request is independent, any healthy target can respond. A service that stores a user’s session only in one server’s memory may require session persistence, often called stickiness, so related requests return to the same target. Stickiness can create uneven loads and complications when that target fails. Many scalable designs instead place session state in a shared store or encode limited state in a protected token. Long-lived connections, retries, timeouts, connection draining, and WebSocket support likewise influence how traffic moves during maintenance or failure.
A load balancer improves availability only when the surrounding design avoids shared points of failure. The balancer itself may need redundant nodes, multiple failure zones, monitored capacity, and resilient name resolution. It cannot make broken application code healthy, prevent overload when every target is saturated, or guarantee that a retry is safe. Logs and metrics are needed to distinguish client errors, balancer limits, and back-end failures. Used well, load balancing separates the stable service entry point from changing compute capacity, allowing traffic to follow healthy resources while maintenance and scaling happen behind that boundary. Traffic distribution also affects deployments. Teams can register a new version with a small weight, observe errors and latency, and increase its share only when results are healthy. During removal, connection draining lets existing requests finish while new requests go elsewhere. This reduces disruption, but it requires timeouts that match real request lengths. Capacity planning still matters because rerouting protects users only when another healthy target has enough resources to accept the additional load.
It applies a configured method such as round robin, least connections, weighting, or a consistent hash among eligible healthy targets.
It tests a target at intervals and temporarily removes that target from new traffic after configured failure conditions are met.
No. It helps route around some failures, but the balancer, application, dependencies, capacity, and network must all be designed for resilience.
Explore more "Explainers"
Discover additional explainers across politics, science, business, technology, and other fields. Each explainer breaks down a complex idea into clear, everyday language—helping you better understand how major concepts, systems, and debates shape the world around us.
