Software

What Is a Load Balancer? How to Set Up Nginx Load Balancing

Talha Aslan 19 min read 2 views

What is a load balancer and what does it actually do?

A load balancer is a layer that sits in front of your servers and spreads incoming requests across several backends instead of sending everything to one machine. Put simply, its job is this: no single server gets overwhelmed, traffic flows around a failed server, and visitors never notice the switch.

Think of a busy bank with a single teller. The queue grows, and then if that teller calls in sick, everything stops. Now picture a host at the door who sends each customer to the next free window. As a result, waiting times drop, and one closed window no longer halts the branch. A load balancer plays that host role for web traffic.

In this guide we explain how load balancing works, the difference between L4 and L7, the algorithms listed in the official Nginx documentation, and a working upstream configuration. We are Talha Aslan and team, a digital marketing and web team, not a hosting company. That is why we base every technical detail on the Nginx docs and skip any default we cannot verify.

Why would a website need load balancing?

The benefit is not only speed. The real gain is that your site stops depending on a single machine. For online stores that see sudden spikes during campaigns, that difference shows up directly in revenue.

  • Traffic distribution: Requests spread across several servers, so no single CPU or memory pool becomes the bottleneck.
  • High availability: If a backend stops responding, the load balancer takes it out of rotation and sends traffic to the others.
  • Zero-downtime maintenance: You pull one server out of the pool to update it while the rest keep serving visitors.
  • Horizontal scaling: Instead of buying one bigger machine, you add another server to the pool.
  • A single entry point: TLS certificates, header rules and access logs live in one place.

That said, load balancing will not fix a slow application on its own. If a page takes three seconds because of a heavy database query, running that query on three servers does not make it faster for each visitor. So we recommend you first rule out the server-side causes of a slow website. Then make the scaling call.

How is a load balancer different from a reverse proxy?

People mix these up because Nginx does both with the same software. A reverse proxy stands between the client and the backend and forwards requests to it. A load balancer also forwards requests, but it splits them across several targets according to a rule.

In other words, every load balancer behaves like a reverse proxy in practice, but not every reverse proxy balances load. If you put Nginx in front of a single Node.js app to handle TLS and caching, that is a reverse proxy setup. Once you run three copies of the same app and let Nginx share requests among them, you are doing load balancing.

We cover the basic Nginx installation and single-target reverse proxy settings in separate articles in this series, so we will not repeat those steps here. This guide assumes Nginx already runs on your server. Our focus is how you manage several servers with the upstream block.

What is the difference between an L4 and an L7 load balancer?

Load balancers fall into two groups based on which network layer they inspect. An L4 (transport layer) balancer only sees connection details such as IP address and port. An L7 (application layer) balancer reads the HTTP request itself, so it can decide based on the path, headers, cookies and hostname.

AspectL4 load balancingL7 load balancing
What it seesIP address, port, protocol (TCP/UDP)HTTP path, headers, cookies, hostname
Where it lives in Nginxstream blockhttp block
Typical useDatabase proxies, DNS, game or custom TCP servicesWebsites, APIs, online stores
Path-based routingNoYes (for example, /api to its own pool)
TLS terminationUsually passes encrypted traffic throughDecrypts here and can read the content
Processing costLowerHigher, but far more flexible

For websites, L7 is usually the right choice because you can route by content. The official Nginx TCP and UDP load balancing guide notes that L4 runs in a separate stream context. In open source Nginx, that module has to be enabled at build time or loaded as a dynamic module. Check whether your distribution package already includes it.

Which load balancing methods does Nginx support?

The official Nginx load balancing page and the upstream module reference list the methods available in the open source build. If you do not specify any method, Nginx uses round robin by default.

  1. Round robin (default): Sends requests to each server in turn. You can also give a stronger server a bigger share with the weight parameter.
  2. least_conn: Sends each new request to the server with the fewest active connections. It balances better when request times vary.
  3. ip_hash: Builds a key from the client IP address, so the same client keeps landing on the same server while it stays up.
  4. hash: You choose the key, for example the request URI. Adding consistent means fewer keys move when you add or remove servers.
  5. random: Picks a server at random. With two, it first picks two candidates and then takes the less loaded one.

The docs also mention least_time, which looks at response time. However, it belongs to the commercial NGINX Plus subscription. So if you add it to an open source config, the configuration test will fail. In short, check each directive against the official page for your version before you rely on it.

Which algorithm should you choose?

The right method depends on how your app behaves. If your servers have similar hardware and requests finish quickly, round robin is usually enough. It is also the simplest option and rarely surprises you.

On the other hand, if request times differ a lot, least_conn gives a fairer split. For example, some requests may return instantly while others build a report for several seconds. Because least_conn looks at current load rather than order, it handles that mix better. File uploads, long API calls and WebSocket connections fit this pattern.

ip_hash is a quick fix for older apps that keep session data on the server's own disk or memory. Still, many users can share a single IP address behind an office network, a mobile carrier or a corporate proxy. In that case one server gets hammered while the others sit idle.

In addition, you will often see hash in front of cache servers. When the same URL always reaches the same cache node, the hit rate goes up. Finally, random two makes sense when several load balancers share the same backend pool. The docs point to distributed setups as its main use case.

How do you set up a load balancer with Nginx?

Load balancing in Nginx has two parts: an upstream block that defines the server pool and a proxy_pass line that sends requests to it. The example below uses three app servers. The IP addresses come from the 192.0.2.x range reserved for documentation, so replace them with your own.

upstream app_backend {
    least_conn;
    server 192.0.2.11:8080 weight=2;
    server 192.0.2.12:8080;
    server 192.0.2.13:8080 backup;
}

server {
    listen 80;
    server_name example.com;

    location / {
        proxy_pass http://app_backend;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}

You typically place this file under /etc/nginx/conf.d/ or your distribution's site directory. Next, run sudo nginx -t to test the syntax. If it passes, apply the change with sudo systemctl reload nginx. A reload picks up the new settings without dropping existing connections.

The header lines also matter. Without them, your backend sees every request as if it came from the load balancer's IP address. That breaks logs, security rules and any IP-based analysis you run later.

What do weight, backup and down do?

Each server line in the upstream block accepts parameters that change its behavior. According to the official upstream reference, the ones you will use most often work like this:

  • weight: Sets the server's share of traffic. The default is 1. In the docs example, a server with weight 3 receives three of every five new requests.
  • backup: Marks the server as a backup. Traffic only moves to it when all primary servers are unavailable.
  • down: Marks the server as permanently unavailable. It is the cleanest way to pull a server out before maintenance.
  • max_conns: Limits simultaneous connections to the server. The default is 0, which means no limit.

There is also one important constraint. According to the docs, you cannot combine backup with the hash, ip_hash or random methods. Therefore, if you plan a backup server, stick with round robin or least_conn.

In practice, a maintenance window looks like this. First, add down to the server you want to update, test the config and reload. Next, wait for active requests on that server to finish and apply your update. After that, remove down and reload again. From the visitor's side, the site never went offline.

How does Nginx handle health checks?

Open source Nginx runs passive health checks. It does not send separate test requests to the backend; instead, it watches the outcome of real visitor requests. If enough attempts fail within a set window, it takes the server out of rotation for a while.

In practice, two parameters control this behavior. max_fails sets how many failed attempts within the window mark a server as unavailable. The default is 1, and a value of 0 turns the counting off. fail_timeout sets both the window for counting failures and how long the server stays out. Its default, in turn, is 10 seconds; after that, Nginx tries the server again.

upstream app_backend {
    server 192.0.2.11:8080 max_fails=3 fail_timeout=30s;
    server 192.0.2.12:8080 max_fails=3 fail_timeout=30s;
}

So what counts as a failed attempt? The proxy_next_upstream directive defines it. By default, connection errors and timeouts count. You can add specific HTTP status codes too, for example with proxy_next_upstream error timeout http_502 http_503;.

One detail deserves attention. By default, Nginx does not retry non-idempotent requests such as POST on another server once the request has reached a backend. That way a checkout form never runs twice. Leave this safeguard alone unless you fully understand the consequences.

Why are active health checks a separate topic?

Passive checks have a weak spot. Nginx only learns that a server is broken after a real visitor lands on it and gets an error. In other words, at least one user sees the failure. Active checks send test requests to every server at regular intervals, independent of traffic, and catch the problem before users do.

The official NGINX HTTP health check guide draws this line clearly. Passive checks exist in both open source Nginx and NGINX Plus. Active checks, which use the health_check directive, are exclusive to commercial NGINX Plus. The guide also says active checks require a shared memory zone in the upstream group.

On the open source side, you have a few ways to close that gap. You can add a simple /health endpoint to your app and poll it with an external monitoring service. Alternatively, you can look at another load balancer with built-in active checks, such as HAProxy, or move to your cloud provider's managed load balancer. The right answer depends on your budget and how much your team can operate.

What is the sticky session problem?

Many apps store a user's login state or cart in the server's own memory or disk. On a single server, that works fine. Once a load balancer joins the picture, however, the first request may hit server A and the second may hit server B. Server B does not know that session, so the user suddenly looks logged out or the cart empties.

There are two ways to deal with this. The first is session affinity, also called sticky sessions: the load balancer keeps sending the same user to the same server. In open source Nginx, the best-known option is ip_hash. The official upstream reference says the cookie-based sticky directive was long limited to the commercial subscription. According to a note on that page, it is part of the open source build from version 1.29.6 onward. Check your version with nginx -v before you count on it.

That said, stickiness has a cost too. When a server crashes, the sessions tied to it disappear anyway. Moreover, load can drift out of balance depending on how users are distributed. That is why we treat stickiness as a transition measure, not a permanent fix.

Why is sharing sessions in Redis the sturdier fix?

The second approach moves sessions off the app servers entirely and into a shared store. Then, whichever server receives the request, it reads the session from the same place. Your backends become stateless, so you can stop or start any of them without affecting users.

Redis is the most common choice for that shared store because it runs in memory and responds very fast. In frameworks like Laravel, switching the session driver to Redis is a configuration change. On the PHP side, you can also point the session handler at Redis with the right extension. We cover Redis for caching and sessions in our Redis vs Memcached caching guide.

Still, keep in mind that a single Redis instance becomes the new single point of failure. For critical projects, plan redundancy for Redis as well. The same logic applies to uploads. If a user's image only lives on server A's disk, server B cannot show it. So you need to write uploads to shared storage or an object store.

How does TLS termination work on a load balancer?

TLS termination, often still called SSL termination, means the load balancer decrypts the HTTPS connection and forwards the request to the backend over the internal network. As a result, you manage the certificate in one place instead of installing it on every app server. Our SSL certificate guide covers the basics.

server {
    listen 443 ssl;
    server_name example.com;

    ssl_certificate     /etc/nginx/ssl/example.com.crt;
    ssl_certificate_key /etc/nginx/ssl/example.com.key;

    location / {
        proxy_pass http://app_backend;
        proxy_set_header Host $host;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}

The X-Forwarded-Proto header plays a key role here. After termination, the backend receives plain HTTP. Without this header, it cannot tell the user arrived over HTTPS and may fall into a redirect loop. In your framework's trusted proxy setting, trust only the load balancer's IP address.

If the network between the load balancer and the backends is shared or untrusted, encrypt that internal traffic too. In that case, use https inside proxy_pass. After setup, check the certificate chain with our SSL checker.

What happens when the load balancer itself fails?

Here is the part many teams miss. Even with three backend servers, a single load balancer in front of them takes the whole site down when it fails. You did not remove the single point of failure; you only moved it.

The classic fix is to run two load balancers in an active and passive pair. Both machines manage a shared virtual IP address through the VRRP protocol, and keepalived is a common open source tool for this. If the active node goes down, the virtual IP moves to the standby and traffic resumes after a short interruption. This setup needs both machines on the same network segment, so not every VPS provider supports it. Ask your provider first.

At larger scale, DNS and anycast methods come into play; we explain the idea in our anycast DNS article. After any change, confirm that your domain resolves to the right IP with our DNS lookup tool. The general name for this mindset is N+1 redundancy: you add at least one spare component on top of the capacity you need. We look at that concept at data center scale in a separate article.

How do cloud load balancers compare with Nginx?

The major cloud providers offer managed load balancing services. With these, the provider handles the load balancer's own redundancy, active health checks and often certificate renewal. You then only define the rules.

Running your own Nginx load balancer gives you full control and predictable costs. You see every line of the config, add the modules you want and avoid lock-in to one provider. The downside is that updates, monitoring and redundancy are entirely on you.

Ask yourself a few questions before you choose. Does anyone on your team respond to a failure at 3 a.m.? Also, does your traffic spike suddenly? And does your infrastructure already live with one cloud provider? If most answers are yes, a managed service usually makes more sense. Our VPS vs cloud server vs VDS comparison helps with the infrastructure side of that call. If you run containers, read our Kubernetes vs Docker guide as well, since Kubernetes distributes traffic in its own service layer.

When is a single server enough?

So let us be honest. Most company websites, blogs and small online stores run comfortably on one well-configured server. Adding a load balancer increases server count, configuration complexity and monthly cost. You should take on that burden only when there is a real need.

A single server is usually enough in these situations:

  • CPU and memory usage stay at a comfortable level even during peak hours.
  • A short planned maintenance window does not seriously hurt your business.
  • The bottleneck is not the server but unoptimized queries, large images or missing caching.
  • Your app keeps sessions and files on local disk and you have no development budget to change that.

For example, if your page speed is poor, caching, image optimization and code fixes deliver a much cheaper win first. Our guide to choosing web hosting also helps you estimate resource needs. A load balancer belongs on the table once those steps run out or uptime becomes truly critical for the business.

When should you leave this setup to someone else?

Setting up load balancing with Nginx is not technically hard. Instead, the hard part is keeping it secure and current for years. On shared hosting you have no access to the Nginx config anyway, and your hosting provider manages it. With managed VPS plans, you should also talk to the provider before you change the infrastructure.

We recommend you hand this job to your hosting provider or an experienced sysadmin when:

  • Nobody on your team has solid experience with Linux server administration and monitoring.
  • The site takes payments and even a few minutes of downtime causes real losses.
  • You also need app changes such as database replication, shared file storage and session migration.
  • Firewalls, DDoS protection and log management all need planning at the same time.

On the web side, we make sure the application is ready for load balancing, meaning sessions and files work in a stateless way. Running the infrastructure should stay with your provider or your systems team. In our custom software development projects, we draw that line from day one.

What should you test after setup?

Reloading the config and opening the site in a browser is not a real test. A load balancer proves its value when something goes wrong. So you should simulate failure on purpose before going live.

  1. Add a temporary response header that differs per backend and confirm that requests really spread out.
  2. Stop the app on one backend. Confirm the site keeps loading and that the Nginx error log shows the matching warning.
  3. Log in, add a product to the cart and browse a few pages. Make sure the session survives.
  4. Check the backend logs and confirm they show the visitor's real IP address.
  5. Test that there is no redirect loop between HTTPS requests and HTTP responses.
  6. Restart the server you stopped and watch it rejoin the pool after the fail_timeout period.

We suggest turning these steps into a checklist and repeating it after every major update. Also include your load balancer config files in your website backup strategy. When a server dies, those are the first files you will need.

What are the most common load balancer mistakes?

Most problems with load balancing do not come from Nginx. Instead, they come from apps that are not ready to run on several servers. That is why you should review the application side as carefully as the config.

The first common mistake is leaving sessions and files on local disk. In practice, the result is users who randomly log out and images that vanish. The second is forgetting the X-Forwarded-For and X-Forwarded-Proto headers. Then analytics and security rules see the wrong IP, and HTTPS redirects loop. The third is running scheduled jobs on every server. For instance, if the same newsletter job fires on three servers, users get three emails.

Another mistake is deploying an older code version to a new server. Automate deployments so every server runs the same release. Finally, teams often forget to monitor the load balancer itself. Watching the backends while ignoring the single machine in front of them brings back the single point of failure we described above.

How should you make the load balancing decision?

To sum up, a load balancer is a layer that spreads traffic to add capacity and hides single-server failures from users. Open source Nginx offers round robin, least_conn, ip_hash, hash and random, along with passive health checks. Active health checks remain an NGINX Plus feature.

Above all, the order of decisions matters. First, measure and optimize your single server. Then make your app stateless by moving sessions to a shared store like Redis and files to shared storage. Only after that should you add a load balancer, and plan a backup for it as well.

If you want your site's infrastructure and marketing goals to move in the same direction, we can work through these questions together in a web design and development project. We get the application ready for load balancing, and we plan infrastructure operations together with your hosting provider.

Frequently Asked Questions

Does a load balancer make a website faster?
Indirectly, yes. A load balancer spreads requests across several servers, which keeps them from overloading and slowing down during peak hours. However, it does not shorten the processing time of a single request. If a page is slow because of a heavy query or large images, fix those first. Otherwise you simply copy the same slowness onto more servers.
Can open source Nginx run active health checks?
No. According to the official NGINX docs, active health checks with the health_check directive are only available in commercial NGINX Plus. Open source Nginx runs passive checks through the max_fails and fail_timeout parameters, which watch the outcome of real requests. For active checks, consider an external monitoring service, another load balancer such as HAProxy, or a managed cloud load balancer.
What is the difference between round robin and least_conn?
Round robin sends requests to servers in turn and is the default method in Nginx. least_conn sends each new request to the server with the fewest active connections. If your requests are short and similar in length, round robin is enough. When some requests run long, such as file uploads or report generation, least_conn gives a more even distribution.
Why do users get logged out behind a load balancer?
If your app stores sessions in a server's own memory or disk, the next request may reach a different server that does not know that session. A quick workaround is a sticky method such as ip_hash. The lasting fix is to move sessions into a shared store like Redis that every server can reach. That makes your app stateless.
Does a small website need a load balancer?
In most cases, no. Company websites, blogs and small online stores run comfortably on one well-configured server. A load balancer adds servers, cost and operational work. It starts to make sense when your resources stay near their limits most of the time, or when even a short outage causes a serious loss of revenue for your business.
  • load balancer
  • load balancing
  • Nginx
  • upstream
  • high availability
  • Redis sessions
  • TLS termination
Share:
Talha Aslan

Google Partner digital marketing expert. Hands-on with SEO, Google Ads, web design and e-commerce projects since 2012; every post here comes from that experience.

Next project

Let's talk about your project.

Your brief goes straight to Talha Aslan and team: strategy led by Talha, delivery by an experienced team. The first consultation is free; we listen and come back with a clear roadmap.