News · New this week

Big Billion Days sale system design: cache, scale, queue

A typical sale-day design reads from cache, scales servers with load, and puts orders in a queue before they reach the database.

The interview question is simple. Why doesn't a Big Billion Days style sale crash at midnight? The answer is a pipeline where each stage takes pressure off the next one.

The path in the reel is this. Phone, then CDN, then app servers, then an order queue, then workers, then the database. Millions of phones hit it at 12 AM, and the database still gets its orders at a steady pace.

Serve product pages from a copy

Start with the CDN. It's a copy kept close to the user. A product page looks the same to everyone, so there's no reason to rebuild it for each request. The CDN serves the cached page and the app servers never see that read.

That's the first cut in load. Anything identical for everyone can come from the copy.

Let app servers scale with load

Some things can't be cached. The cart is the obvious one, since it's different for every user. That traffic goes to the app servers.

This is where autoscaling comes in. When load goes up, new servers start on their own. When it drops, they shut down. In the reel's diagram the app tier runs from 10 to 200 servers, so it can grow for the spike and shrink after it.

Put a queue in front of the database

Scaling the app tier creates a new problem. If every order goes straight to the database, the spike lands there all at once. The user sees a spinning screen.

The fix is an order queue, Kafka in the diagram. It holds the spike and turns it into a line. Orders wait their turn instead of all arriving together.

Workers sit at the other end of that line. They pick up one order at a time, at a steady speed, and write it to the database. The database sees a steady rate, even though the front door took a huge burst.

The database holds stock and orders. Stock goes down as each order is written, and the last phone is sold once, not twice.

The one-line version

Read from cache, buy through a queue, and let servers grow with the load. If you can say that and draw the six boxes, you've covered the core of the answer.

The reel ends with a question worth thinking about. Would you queue users or just add more servers? Before your next system design interview, sketch this pipeline from memory and explain what breaks if you remove any one stage.

  • #systemdesign
  • #autoscaling
  • #caching
  • #messagequeue

More reels

All news →