SGLang · LLM Infrastructure

The SGLang System Design Interview

Preparation for the LLM inference infrastructure design interview — the one where you're asked why p99 time-to-first-token went to eight seconds after a change that improved average throughput, and expected to reason about memory bandwidth in the answer.

51Sections
102,879Words
186Drills
L5–L7Ladder
Book One

Deep Dossiers

Thirty-four full interview scenarios with capacity math, failure analysis, and an L5→L6→L7 follow-up ladder.

I — Core Inference ServingII — Long Context & MultimodalIII — Agentic & Structured OutputIV — Reinforcement LearningV — Training SystemsVI — Reliability & Economics
39 sections · 90,161 wordsOpen →
Book Two

Rapid Drills

186 short questions for finding gaps and building recall speed.

Fundamentals & the RooflineKV Cache & MemoryBatching & SchedulingPrefix Caching & RoutingParallelismMixture of ExpertsDisaggregation (PD & EPD)Quantization & KernelsSpeculative & Structured DecodingLong Context & MultimodalRL & TrainingOperations, Cost & Incidents
12 sections · 12,718 wordsOpen →

How to actually practice

Reading this cover to cover will make you feel prepared and will not make you prepared. The failure mode is recognition without recall.

  1. Read only the prompt. Cover the rest. Set a timer for forty minutes.
  2. Work it out loud. On paper, as if the interviewer were there. The common real-interview failure isn't bad thinking — it's thinking silently and presenting a conclusion nobody could follow.
  3. Do the arithmetic. Not “we'd need a lot of GPUs” — the number. Being wrong by 2x is fine; refusing to compute is what gets marked down.
  4. Then read the dossier. Diff your answer against it. Pay attention to what you didn't think to ask, not just what you got wrong.
  5. Come back in a week. Redo it from the prompt alone. Recognition isn't recall.