Skip to content

[Problem] Prefill/Decode (P/D) Disaggregation in v3/llm-inference #32

Description

@weichenxu923

Problem title

Prefill/Decode (P/D) Disaggregation — v3/llm-inference/pd-disaggregation

Why is this worth practising? Who asks it?

Prefill and decode have very different compute profiles. Prefill is compute-bound: one large parallel pass over the prompt. Decode is memory-bandwidth-bound: one token per step, dominated by KV-cache reads. When both phases share a GPU, long prefills stall in-flight decodes and hurt time-between-tokens (TBT), while decode batches waste the compute that prefill could use.

Production serving stacks now run the two phases on separate worker pools and ship the KV cache between them. Examples include DistServe, Splitwise, Mooncake (Kimi), vLLM's disaggregated prefill, SGLang PD, and NVIDIA Dynamo. Inference and infra roles ask about this regularly, e.g. "how would you serve a long-context model with strict TTFT and TBT SLOs?" Likely companies include Together AI, Fireworks AI, Perplexity, Anyscale, and the frontier labs.

It would be the natural next step after the current v3/llm-inference sequence (kv-cache → continuous-batching → inference-engine). Today inference-engine runs prefill and decode inside one engine, and nothing in the repo covers splitting them.

Difficulty

expert

Proposed problem sketch

Build it on the MiniTransformer / KVCache from inference-engine, CPU-only, and simulate the workers in one process:

  1. PrefillWorker: takes a batch of prompts, runs one forward pass, and returns the first sampled token plus a per-request KV cache.
  2. KVTransfer: serializes and hands off each request's KV cache (optionally as fixed-size blocks/pages) from the prefill worker to the decode worker, and tracks transferred bytes.
  3. DecodeWorker: keeps a continuous batch of running requests, admits new requests as their KV arrives, runs one decode step per tick, and evicts finished requests.
  4. Scheduler / router: sends new requests to prefill, sends completed prefills to decode, and reports TTFT and per-token latency per request.

Checks (auto-gradable):

  • Tokens generated by the disaggregated pipeline exactly match a colocated baseline (InferenceEngine) under greedy decoding.
  • The KV cache received by decode is bitwise equal to what prefill produced (shape (B, heads, seq_len, d_head) per layer).
  • Decode keeps making progress (no stalled ticks) while a long prompt is being prefilled, which is the core property colocated scheduling lacks.
  • The KV transfer size matches 2 * num_layers * seq_len * d_model * dtype_bytes per request.

Possible follow-ups (interview-style): choosing the prefill:decode worker ratio, chunked prefill as the alternative to disaggregation, KV transfer cost vs. prompt length, and prefix-cache reuse on the prefill side.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions