Preparation for the LLM inference infrastructure design interview — the one where you're asked why p99 time-to-first-token went to eight seconds after a change that improved average throughput, and expected to reason about memory bandwidth in the answer.
Thirty-four full interview scenarios with capacity math, failure analysis, and an L5→L6→L7 follow-up ladder.
186 short questions for finding gaps and building recall speed.
Reading this cover to cover will make you feel prepared and will not make you prepared. The failure mode is recognition without recall.