Load balancing
What is load balancing?
Load balancing is the practice of distributing incoming network traffic across multiple servers, so that no single server gets overwhelmed while others sit idle. A load balancer sits in front of a pool of servers and decides, for each incoming request, which server should handle it, based on an algorithm such as round robin or least connections. This is what lets a website or an API keep serving traffic reliably as usage grows, and keep working even if one server in the pool fails.
Without load balancing, scaling a service means making one server bigger (vertical scaling), which has hard physical limits; with load balancing, a team can add more servers instead (horizontal scaling) and let the load balancer spread traffic between them.
How a load balancer works
A load balancer typically also runs health checks against each server in the pool, and stops routing traffic to any server that fails to respond, so users never hit a dead instance. A minimal health check configuration for a load balancer might look like this:
upstream backend {
server 10.0.0.1:8080 max_fails=3 fail_timeout=30s;
server 10.0.0.2:8080 max_fails=3 fail_timeout=30s;
server 10.0.0.3:8080 max_fails=3 fail_timeout=30s;
}
Load balancers operate at different layers of the network: a Layer 4 load balancer routes traffic based on IP address and port, while a Layer 7 load balancer can inspect the actual HTTP request (URL path, headers, cookies) and route more intelligently, for example sending all API traffic to one pool of servers and all static asset requests to another.
Common load balancing algorithms
| Algorithm | How it distributes traffic | Best suited for |
|---|---|---|
| Round robin | Sends each new request to the next server in sequence | Servers with similar capacity |
| Least connections | Sends requests to the server with the fewest active connections | Long-lived or uneven-duration connections |
| IP hash | Routes a given client IP to the same server consistently | Session data stored locally on a server |
| Weighted round robin | Distributes traffic proportionally to server capacity | Servers with different hardware specs |
Types of load balancers
- Hardware load balancer: a dedicated physical appliance, common in large on-premise data centers.
- Software load balancer: a program such as Nginx or HAProxy running on a standard server or virtual machine.
- Cloud load balancer: a managed service offered by a cloud or CDN provider, scaling automatically without infrastructure to maintain.
- DNS load balancing: distributes traffic by returning different server IPs to different clients at the DNS level, a coarser but simple method.
Best practices and common pitfalls
- Always pair load balancing with health checks, so traffic automatically stops going to a failed or slow server instead of degrading the experience for real users.
- Choose sticky sessions (routing a user consistently to the same server) only when necessary, since they reduce the flexibility of the load balancer and can create hot spots.
- Monitor per-server response times and error rates, not just overall traffic, to catch one underperforming server before it drags down the whole pool.
- Plan for the load balancer itself to be redundant; a single load balancer with no failover becomes a new single point of failure.
Why load balancing matters for reliability
Load balancing is what makes horizontal scaling possible, letting infrastructure grow by adding servers rather than upgrading one machine indefinitely, and it is a core building block of high-availability architecture: if one server crashes or needs maintenance, the load balancer simply routes around it, and users notice nothing. For any product expecting real growth in traffic, load balancing is one of the first pieces of infrastructure to get right, well before it becomes a bottleneck.
Load balancing at BeBranded
When we architect hosting and infrastructure for clients, we put load balancing in place as soon as reliability or expected traffic justifies more than a single server, whether through a cloud provider's managed load balancer or a CDN edge network. This keeps a site or an app responsive during traffic spikes and resilient to a single server going down. See our website service for how we set up hosting and infrastructure.