System Design Interview Questions: What 2026 Loops Test

System design interviews in 2026 reward numbers, failure modes and cost, not boxes and arrows. Here are the five areas every loop probes, plus a 6-question self-check to find your gaps first.

18 min read

System design interviews got harder in 2026 without getting longer. More candidates can draw the standard picture of load balancer, service, cache and database than there are senior roles to fill, so the picture stopped separating anyone. What interviewers listen for now is what happens when the design meets a number: a load estimate, a hot key during a flash sale, a shard count that has to double, a payment request that arrives twice.

That changes what preparing for system design interview questions should look like. Memorising twenty case studies is the slow way. The follow-up questions that decide the round come from a small set of mechanisms, and almost all of them sit in five areas:

  1. Requirements and estimation - turning "fast and scalable" into numbers a design can be checked against
  2. Caching - where stale data, stampedes and hot keys come from
  3. Partitioning and sharding - what moves when you add a machine, and what gets hot when you do not
  4. Consistency and coordination - CAP, replication lag and deciding which node is in charge
  5. Asynchronous messaging and retries - queues, logs, and making a repeated request harmless

The sections below go through them in that order. If you only have an evening, take the quiz first and skip what you already get right.

Start here: a 6-question self-check

These six come from the system design set on squizzu: one for each area, and a second one on retries because that is where designs quietly break. Pick an answer and you get the reasoning straight away, plus a longer write-up if you want it. You do not need an account.

Squizzu Logo
System Design • Interview Framework & Estimation

Question 1 / 6

By Little's Law, how many requests are in flight on average in a service that handles 2,000 requests per second with an average latency of 50 ms?

The hot key and the duplicate shipment are the two that people with years of backend work still get wrong. The hot key catches people because the instinctive fix of adding cache nodes changes nothing at all. The shipment race is harder to spot: the code that causes it looks correct in review and only fails under concurrent retries. If either one surprised you, start with areas 2 and 5.

What a system design interview looks like in 2026

The format has barely changed: one open prompt, 45 to 60 minutes, a whiteboard or a shared diagramming tool. You will be asked to design something familiar, such as a URL shortener, a news feed, a chat service, a rate limiter or a ride-matching system, and the interviewer will steer the second half of the conversation towards the parts they find interesting.

What changed is the scoring. Three shifts show up across companies:

  • Numbers over boxes. A design that is never sized cannot be wrong, which is exactly why interviewers now push for estimates early and then hold you to them.
  • Failure modes and cost. "What happens when this node dies?" and "What does this cost per month?" used to be senior-only follow-ups. They are now standard, partly because cloud bills became an engineering concern.
  • LLM-backed prompts. "Design the serving layer for a chat assistant" or "add semantic search to this product" now appear alongside the classics. They do not replace the five areas below. They add a very expensive, very slow dependency to them: GPU capacity to plan, token streaming to deliver, and responses worth caching when prompts repeat. If the role is AI-heavy, how to prepare for an AI engineer interview covers the model side.

The time budget is the constraint most candidates underestimate. With five minutes of introductions and five of questions at the end, you have about 40 minutes. A workable split is five for requirements, five for estimation, ten for a high-level design and twenty for one or two deep dives.

Area 1: Requirements and back-of-the-envelope estimation

Every design discussion starts with requirements, and most candidates treat this step as a formality. That is the first mistake interviewers notice.

Functional requirements say what the system does: users post photos, followers see them in a feed. Non-functional requirements say how well it has to do it: latency, availability, durability, consistency. The functional list is usually easy. The non-functional one is where the design gets decided, and only if it is written as numbers. "The feed should be fast" cannot rule out any design. "The feed loads in under 200 ms at p99 for 100 million daily users" rules out computing the feed from scratch on every request, which is the first real decision in the problem.

Estimation turns those numbers into sizes. The interviewer is not checking arithmetic to three significant figures. They are checking whether you know which numbers matter and roughly how big they are. The set worth having ready:

Quantity Rule of thumb
Seconds in a day about 86,400, so 1 million requests a day is about 12 per second
Peak vs average plan for 2 to 3 times the daily average unless told otherwise
Round trip within a data center about 0.5 ms
Round trip across a continent and back about 150 ms
Random read from SSD on the order of 100 microseconds
Cache sizing the hot 20% of a day's distinct items is a common starting point

A little queueing theory comes up more often than candidates expect. You only need two results.

Little's Law says the average number of requests inside a system equals the arrival rate multiplied by the average time each one spends there. It sizes thread pools, connection pools and in-flight limits, and it needs nothing except consistent units. The usual slip is multiplying requests per second by a latency still written in milliseconds.

Utilisation is not linear. In the simplest queueing model, a single server with random arrivals, the mean response time is the service time divided by one minus the utilisation. That formula looks harmless and behaves very badly near the top:

Line chart of mean response time against server utilisation for a single-server queue. Response time is twice the service time at 50% utilisation, five times at 80% and ten times at 90%, rising steeply towards the right edge of the chart.

Going from 50% to 90% busy does not make requests 1.8 times slower. It makes them five times slower. That is why capacity plans target 60 to 70% utilisation, and it is the reason "the servers are only at 85% CPU" is not a reassuring sentence in an incident review.

What the interviewer is testing: whether your design is driven by the requirements or merely decorated with them. Common follow-up: "Why did you choose this database?" The answer they want names the requirement it serves and what the choice gives up. Naming the companies that use it is not an answer.

Area 2: Caching, and the three ways it goes wrong

Caching is the first tool candidates reach for, and it is where the follow-up questions go deepest, because a cache adds a second copy of the truth and every problem in this area comes from that.

The default pattern is cache-aside: the application checks the cache, and on a miss it reads the database and writes the result into the cache itself. It is simple, and it fails in three predictable ways.

Stale data. With only a TTL, a changed record stays wrong in the cache until the entry expires. The usual fix is to delete the key when the underlying row is written, so the next read misses and loads the new value. Deleting is safer than updating the cache in place, because two concurrent writers updating the cache can leave it holding the older of the two values.

Stampedes. A popular key expires, hundreds of concurrent requests miss at the same moment, and all of them run the expensive query that rebuilds it. The database sees a burst it was never sized for, exactly when the cache was supposed to protect it. The remedies are to let only one request rebuild the value while the others wait or receive the old copy, to refresh popular keys shortly before they expire, and to add random jitter to TTLs so keys written together do not expire together.

Hot keys. A distributed cache assigns each key to one node. That is how it scales, and it is also why one extremely popular key, such as a flash-sale product, lands all its reads on a single machine. The node saturates while the rest of the cluster idles, and adding nodes changes nothing, because the key still lives in one place. The fix is to stop sending every read of that key to the cluster: serve it from a small in-process cache on each application server with a very short TTL, or replicate it under several suffixed keys so reads spread out. Both trade a few seconds of staleness for a load profile the cluster can survive, so check that the product can accept it.

A few facts worth having straight because interviewers like them as short checks:

  • Redis vs Memcached. Redis offers server-side data structures such as hashes, sorted sets and lists; Memcached stores opaque values. Both support TTLs and LRU-style eviction.
  • Redis Cluster placement. Keys map to 16,384 hash slots through CRC16, and slots are moved between masters when resharding. It is not a consistent-hashing ring, which surprises people.
  • Pull vs push CDN. A pull CDN fetches a file from the origin on the first request at each edge. A push CDN is filled ahead of time.

What the interviewer is testing: that you see the failure a cache introduces, not only the latency it removes. Common follow-up: "What happens to the database when the cache cluster restarts?" A cold cache is a stampede on every key at once, and a good answer mentions warming it or rate-limiting the rebuild.

Area 3: Partitioning, sharding and consistent hashing

Sharding comes up once the data or the write volume no longer fits on one machine, and interviewers probe two things: how keys are placed, and what happens when that placement has to change.

Range vs hash partitioning. Range partitioning keeps keys sorted, so a query over a range of dates or names touches only a few partitions. The price is hot spots: if keys arrive in increasing order, such as timestamps, every new write lands on the last partition. Hash partitioning spreads writes evenly and destroys ordering, so range queries have to ask every partition. Wide-column stores like Cassandra combine the two: the first part of the key is hashed to choose the partition, and a clustering column keeps rows sorted inside it, which is how "one user's posts between two dates" becomes a single contiguous read.

Placement by hash(key) mod N is the obvious scheme and the one interviewers like to break. Changing N changes the remainder for most keys, and most keys therefore move. How many move depends on how N changes, and working that out from the remainder arithmetic is a common interview question. Adding a single shard is close to the worst case.

Consistent hashing exists to fix that. Nodes and keys are hashed onto the same ring, and each key belongs to the next node clockwise. Adding a node takes over only the keys between it and its predecessor, so roughly one Nth of the data moves instead of most of it. The catch interviewers ask about is balance: with one position per node, random placement leaves some arcs of the ring far longer than others, and the busiest node can own several times its intended share. Virtual nodes fix that by giving each physical node many positions on the ring, which averages the arcs out and also lets a larger machine take more positions.

Two panels comparing key movement when a fifth shard is added. With hash mod N, about 80% of keys change shard. With consistent hashing, only the keys on one arc of the ring move, about 20%.

Name these two patterns if the conversation gets there:

  • Write sharding for a hot partition key. If one key, such as today's date, receives all writes, split it into N sub-keys. Compute the suffix from something the reader also knows, such as a hash of the order ID. A random suffix spreads writes just as well but forces every read to query all N sub-keys.
  • Reference tables. A small, rarely changing table that every shard joins against, such as airports or currencies, is copied in full to every shard so the join stays local.

What the interviewer is testing: whether you choose a shard key from the access pattern, not from the schema. Common follow-up: "Which queries does this shard key make expensive?" Every shard key makes some query a scatter-gather across all shards; say which one and why you can live with it.

Area 4: Consistency, replication and coordination

This is the area with the most vocabulary and the fewest candidates who can apply it.

CAP is narrower than its reputation. It only describes what a system does during a network partition: either it refuses some requests so that no replica serves stale data (consistency), or it keeps answering on both sides and lets the replicas diverge until the partition heals (availability). "CA" is not a real option for a distributed system, because partitions are not optional. PACELC adds the part that matters the rest of the time: else, when there is no partition, you trade latency against consistency. Waiting for a remote replica to confirm a write costs a round trip on every request.

The interview question is almost never "define CAP". It is a scenario: a shopping cart that must keep accepting items during a data-center split, or an account balance that must never go negative. The first can accept writes on both sides and merge them later. The second cannot. Being able to say which features of one product need which choice is the whole skill.

Replication lag is the practical face of this. With asynchronous replication, a follower applies the leader's writes some time after they commit, usually milliseconds and occasionally minutes. A user who updates their profile and then reads from a lagging replica sees the old version. The fixes are consistency guarantees scoped to what the user needs: read-your-writes (route a user's reads to the leader for a short time after they write) or monotonic reads (pin a user to one replica). Knowing that eventual consistency only promises that replicas converge once writes stop, with no deadline, keeps you from overselling it.

Coordination is the other half. Some work must be done by exactly one node: a nightly job, a partition's leader, a lock holder. The standard answer is leader election through a lease in a coordination service such as etcd or ZooKeeper. The lease expires if its holder dies, so another node can take over. The reason you need that machinery rather than a simple timeout is worth stating: a timeout cannot tell a crashed node from a slow one. A long garbage-collection pause looks exactly like death, and a node that wakes up still believing it is the leader will act on stale authority unless every write carries a fencing token the storage layer checks.

For workflows that span several services, sagas replace distributed transactions: each step commits locally and a failure triggers compensating steps. The follow-up question is about isolation, because other processes can see a saga's intermediate state. Marking records as pending until the saga completes is the usual countermeasure.

What the interviewer is testing: consistency chosen per feature, not one setting for the whole system. Common follow-up: "What does the user see while the replicas disagree?" Answer that concretely and most of this area follows.

Area 5: Queues, event streams and idempotent retries

Almost every design ends up with something asynchronous, because slow or unreliable work does not belong inside a user's HTTP request. Transcoding a video, sending an email, charging a card: the endpoint records the intent, enqueues a job, returns an ID, and workers do the rest.

The two families of broker behave differently, and interviewers check that you know which one you are drawing:

Message queue (SQS, RabbitMQ) Log (Kafka, Kinesis)
After a message is processed deleted kept until retention expires
Consumer position tracked by the broker an offset each consumer group owns
Replay not possible rewind the offset
Ordering usually none, or per group per partition
Parallelism add workers freely capped by the partition count

Most follow-up questions about logs come from a pair of properties. Order is guaranteed only within a partition, so events that must be applied in order, such as the status changes of one shipment, need a partition key that sends them to the same partition. And within one consumer group, each partition is read by at most one consumer. A group with more consumers than partitions has idle members, and adding more does nothing. That second point is easy to miss in an incident, when the obvious reaction to growing lag is to scale out.

Delivery guarantees are the part that has to be said precisely. At-least-once delivery means nothing is lost but anything can arrive twice: the broker retries until it gets an acknowledgement, and an acknowledgement can be lost after the work was done. Exactly-once processing end to end is something you build, not something you buy. You build it by making the consumer idempotent, so that applying the same message twice has the same effect as applying it once.

The same problem appears at the API edge. A mobile client times out on POST /orders and retries, and the customer gets two orders. The standard fix is an idempotency key: the client generates one value per logical operation and sends it on every retry, and the server returns the stored result when it sees the key again. The details are where candidates lose points. The key has to be created per operation, not per HTTP attempt. On the server, "look up the key, and if it is missing, do the work and then store it" is a race, because two concurrent retries can both pass the lookup. The key must be claimed atomically before the side effect, and a database transaction cannot roll back a call to an external API that has already happened.

What the interviewer is testing: whether you design for the retry you know will happen. Common follow-up: "What happens if the worker crashes halfway through?" Walk through redelivery, the idempotency check and what a half-finished job leaves behind.

A one-week system design interview plan

Written for people who already build backend services. The point is to say things out loud, in order, against a clock.

  • Day 1 - Numbers. Memorise the estimation table above, then size three systems you use daily: requests per second, storage per year, cache memory. Time yourself; five minutes each is the target.
  • Day 2 - Requirements. Take three classic prompts and write only the requirements, quantified. Check that each non-functional number would actually rule out some design.
  • Day 3 - Caching. For one prompt, design the cache layer and then list how it fails: stale reads, a stampede, a hot key, a cold restart. Write one sentence of mitigation for each.
  • Day 4 - Sharding. Pick a shard key for a feed, a chat service and an order system. For each, name the query it makes expensive.
  • Day 5 - Consistency. Split one product into features and assign each a consistency requirement. A cart, a balance, a like counter and a profile page should not all get the same answer.
  • Day 6 - Async. Redesign a slow synchronous endpoint around a queue, then trace a worker crash and a client retry end to end and show where idempotency keeps the result correct.
  • Day 7 - Rehearse. Run one full prompt out loud in 40 minutes, with a timer and ideally another person. The first attempt usually overruns on requirements; that is the thing to fix.

Where to practise next

The self-check only samples the topic. The full system design set on squizzu goes much further, from estimation and caching through consensus, messaging and case studies such as URL shorteners, feeds and booking systems.

Work through the system design questions on squizzu. When you miss one, the explanation walks through every option, so you come away knowing which mechanism to revisit rather than which letter to remember.

If the loop also has a language round, Python interview questions for 2026 covers the one most backend and data teams use.

The storage layer usually gets its own round: the SQL quiz covers the queries and indexes underneath every design above, and the AWS quiz covers the managed services most designs end up running on. If the role is closer to data platforms, data engineer interview questions for 2026 takes partitioning, streams and idempotency into Spark and Kafka. If the prompt is an LLM feature, how to prepare for an AI engineer interview covers retrieval, evaluation and inference cost.

Frequently asked questions

What questions are asked in a system design interview?

Most rounds are one open prompt, such as design a URL shortener, a news feed, a chat system or a rate limiter, followed by follow-up questions that test five areas: turning vague requirements into numbers, caching, partitioning and sharding, consistency and coordination between replicas, and asynchronous messaging with retries. The prompt changes from company to company; the follow-ups barely do. In 2026 a growing share of prompts involve an LLM-backed feature, which adds GPU capacity, token streaming and inference cost to the same five areas rather than replacing them.

How do I prepare for a system design interview in 2026?

Learn the five recurring areas as mechanisms rather than as vocabulary, practise back-of-the-envelope estimation until you can size traffic, storage and cache memory in a couple of minutes, and rehearse a small number of classic prompts out loud under a 45-minute limit. For each component you draw, be ready to say what requirement it serves, what it costs, and how it fails. Two weeks of an hour a day is enough for most engineers who already build backend services.

What is the best framework for a system design interview?

Four steps in order: clarify functional and non-functional requirements and quantify the non-functional ones, estimate load and storage, sketch a high-level design that meets the requirements with the simplest components possible, then go deep on the one or two parts the interviewer cares about. The step candidates skip most often is quantifying requirements, and it is the one that makes the rest of the discussion decidable.

Do I need to know CAP theorem for system design interviews?

Yes, but the useful version is narrower than most summaries. CAP only describes what happens during a network partition: a system either refuses some requests to stay consistent or keeps answering and lets replicas diverge. PACELC adds the trade-off that applies the rest of the time, between latency and consistency. Interviewers rarely ask for the definition. They ask which choice a specific feature needs, such as a shopping cart versus a bank balance, and why.

Are system design interviews required for junior developers?

Usually not as a full round. Most companies start system design interviews at mid-level or senior, and junior loops replace them with an object-oriented design or API design discussion. That said, junior candidates are increasingly asked lighter design questions, such as how they would cache a slow endpoint or why a background job should use a queue, so the fundamentals in this article are worth knowing at any level.

How long is a system design interview?

Typically 45 to 60 minutes, of which roughly ten go to introductions and to your questions at the end. That leaves about 40 minutes for requirements, estimation, a high-level design and one or two deep dives, which is why spending fifteen minutes on requirements or drawing every component in detail is the most common way to run out of time.

We use cookies

Some cookies are needed to run this site. With your consent we also measure how it is used, so that we can improve it.

Cookie policy

Choose what we may measure. You can change this at any time.

Strictly necessary

Essential for the proper functioning of the website. These cannot be disabled.

Performance and analytics

Help us understand how Squizzu is used, diagnose technical issues and improve the service.

System Design Interview Questions: What 2026 Loops Test | Squizzu