Juan Pablo García
All writing
Explainer5 steps · scroll to play

How a web request is served

Most web systems follow the same path: a load balancer spreads requests across servers, a cache answers the common ones fast, and a database holds the truth.

  1. 1

    A request arrives

    A browser sends a request. Before any of your code runs, it reaches a load balancer.

  2. 2

    Spread the load

    The load balancer sends each request to a different server, so no single machine gets overwhelmed.

  3. 3

    Check the cache first

    The server asks an in-memory cache for the answer. A hit comes back in a couple of milliseconds.

  4. 4

    Fall back to the database

    On a miss, the server queries the database, which is slower, then saves the result in the cache for the next request.

  5. 5

    Scale out

    When traffic grows, add servers behind the balancer. Watch the slowest requests (p99), not only the average.

A request arrives

A browser sends a request. Before any of your code runs, it reaches a load balancer.

Spread the load

The load balancer sends each request to a different server, so no single machine gets overwhelmed.

Check the cache first

The server asks an in-memory cache for the answer. A hit comes back in a couple of milliseconds.

Fall back to the database

On a miss, the server queries the database, which is slower, then saves the result in the cache for the next request.

Scale out

When traffic grows, add servers behind the balancer. Watch the slowest requests (p99), not only the average.

In short

  • A cache is the cheapest speedup. Keeping it from serving stale data is the hard part.
  • Stateless servers are easy to scale out. Keep sessions and data in shared stores.
  • The latencies in the animation are illustrative. Measure your own percentiles.

Have an AI feature to build? Let's talk for 15 minutes.