Problem title
Prefill/Decode (P/D) Disaggregation — v3/llm-inference/pd-disaggregation
Why is this worth practising? Who asks it?
Prefill and decode have very different compute profiles. Prefill is compute-bound: one large parallel pass over the prompt. Decode is memory-bandwidth-bound: one token per step, dominated by KV-cache reads. When both phases share a GPU, long prefills stall in-flight decodes and hurt time-between-tokens (TBT), while decode batches waste the compute that prefill could use.
Production serving stacks now run the two phases on separate worker pools and ship the KV cache between them. Examples include DistServe, Splitwise, Mooncake (Kimi), vLLM's disaggregated prefill, SGLang PD, and NVIDIA Dynamo. Inference and infra roles ask about this regularly, e.g. "how would you serve a long-context model with strict TTFT and TBT SLOs?" Likely companies include Together AI, Fireworks AI, Perplexity, Anyscale, and the frontier labs.
It would be the natural next step after the current v3/llm-inference sequence (kv-cache → continuous-batching → inference-engine). Today inference-engine runs prefill and decode inside one engine, and nothing in the repo covers splitting them.
Difficulty
expert
Proposed problem sketch
Build it on the MiniTransformer / KVCache from inference-engine, CPU-only, and simulate the workers in one process:
PrefillWorker: takes a batch of prompts, runs one forward pass, and returns the first sampled token plus a per-request KV cache.
KVTransfer: serializes and hands off each request's KV cache (optionally as fixed-size blocks/pages) from the prefill worker to the decode worker, and tracks transferred bytes.
DecodeWorker: keeps a continuous batch of running requests, admits new requests as their KV arrives, runs one decode step per tick, and evicts finished requests.
Scheduler / router: sends new requests to prefill, sends completed prefills to decode, and reports TTFT and per-token latency per request.
Checks (auto-gradable):
- Tokens generated by the disaggregated pipeline exactly match a colocated baseline (
InferenceEngine) under greedy decoding.
- The KV cache received by decode is bitwise equal to what prefill produced (shape
(B, heads, seq_len, d_head) per layer).
- Decode keeps making progress (no stalled ticks) while a long prompt is being prefilled, which is the core property colocated scheduling lacks.
- The KV transfer size matches
2 * num_layers * seq_len * d_model * dtype_bytes per request.
Possible follow-ups (interview-style): choosing the prefill:decode worker ratio, chunked prefill as the alternative to disaggregation, KV transfer cost vs. prompt length, and prefix-cache reuse on the prefill side.
Problem title
Prefill/Decode (P/D) Disaggregation —
v3/llm-inference/pd-disaggregationWhy is this worth practising? Who asks it?
Prefill and decode have very different compute profiles. Prefill is compute-bound: one large parallel pass over the prompt. Decode is memory-bandwidth-bound: one token per step, dominated by KV-cache reads. When both phases share a GPU, long prefills stall in-flight decodes and hurt time-between-tokens (TBT), while decode batches waste the compute that prefill could use.
Production serving stacks now run the two phases on separate worker pools and ship the KV cache between them. Examples include DistServe, Splitwise, Mooncake (Kimi), vLLM's disaggregated prefill, SGLang PD, and NVIDIA Dynamo. Inference and infra roles ask about this regularly, e.g. "how would you serve a long-context model with strict TTFT and TBT SLOs?" Likely companies include Together AI, Fireworks AI, Perplexity, Anyscale, and the frontier labs.
It would be the natural next step after the current
v3/llm-inferencesequence (kv-cache→continuous-batching→inference-engine). Todayinference-engineruns prefill and decode inside one engine, and nothing in the repo covers splitting them.Difficulty
expert
Proposed problem sketch
Build it on the
MiniTransformer/KVCachefrominference-engine, CPU-only, and simulate the workers in one process:PrefillWorker: takes a batch of prompts, runs one forward pass, and returns the first sampled token plus a per-request KV cache.KVTransfer: serializes and hands off each request's KV cache (optionally as fixed-size blocks/pages) from the prefill worker to the decode worker, and tracks transferred bytes.DecodeWorker: keeps a continuous batch of running requests, admits new requests as their KV arrives, runs one decode step per tick, and evicts finished requests.Scheduler/ router: sends new requests to prefill, sends completed prefills to decode, and reports TTFT and per-token latency per request.Checks (auto-gradable):
InferenceEngine) under greedy decoding.(B, heads, seq_len, d_head)per layer).2 * num_layers * seq_len * d_model * dtype_bytesper request.Possible follow-ups (interview-style): choosing the prefill:decode worker ratio, chunked prefill as the alternative to disaggregation, KV transfer cost vs. prompt length, and prefix-cache reuse on the prefill side.