How a web request is served
Most web systems follow the same path: a load balancer spreads requests across servers, a cache answers the common ones fast, and a database holds the truth.
- 1
A request arrives
A browser sends a request. Before any of your code runs, it reaches a load balancer.
- 2
Spread the load
The load balancer sends each request to a different server, so no single machine gets overwhelmed.
- 3
Check the cache first
The server asks an in-memory cache for the answer. A hit comes back in a couple of milliseconds.
- 4
Fall back to the database
On a miss, the server queries the database, which is slower, then saves the result in the cache for the next request.
- 5
Scale out
When traffic grows, add servers behind the balancer. Watch the slowest requests (p99), not only the average.
A request arrives
A browser sends a request. Before any of your code runs, it reaches a load balancer.
Spread the load
The load balancer sends each request to a different server, so no single machine gets overwhelmed.
Check the cache first
The server asks an in-memory cache for the answer. A hit comes back in a couple of milliseconds.
Fall back to the database
On a miss, the server queries the database, which is slower, then saves the result in the cache for the next request.
Scale out
When traffic grows, add servers behind the balancer. Watch the slowest requests (p99), not only the average.
In short
- A cache is the cheapest speedup. Keeping it from serving stale data is the hard part.
- Stateless servers are easy to scale out. Keep sessions and data in shared stores.
- The latencies in the animation are illustrative. Measure your own percentiles.