[https://nvbugs/6465993][fix] use attention cache dtype for disaggregated transfer - #16505
Conversation
|
/bot run --disable-fail-fast --stage-list "DGX_B200-8_GPUs-PyTorch-1, DGX_B200-8_GPUs-PyTorch-2, DGX_B200-8_GPUs-PyTorch-4" |
|
/bot run --disable-fail-fast --stage-list "DGX_B200-8_GPUs-PyTorch-1, DGX_B200-8_GPUs-PyTorch-2, DGX_B200-8_GPUs-PyTorch-4" |
|
PR_Github #59749 [ run ] triggered by Bot. Commit: |
|
PR_Github #59749 [ run ] completed with state
|
|
/bot run --disable-reuse-test --stage-list "DGX_B200-8_GPUs-PyTorch-4" |
|
PR_Github #59836 [ run ] triggered by Bot. Commit: |
|
PR_Github #59836 [ run ] completed with state
|
|
/bot run --disable-fail-fast --disable-reuse-test --stage-list "DGX_B200-8_GPUs-PyTorch-4" |
|
PR_Github #59852 [ run ] triggered by Bot. Commit: |
|
PR_Github #59852 [ run ] completed with state
|
|
/bot run --disable-fail-fast --disable-reuse-test --stage-list "DGX_B200-8_GPUs-PyTorch-4" |
|
PR_Github #59892 [ run ] triggered by Bot. Commit: |
|
PR_Github #59892 [ run ] completed with state
|
|
/bot run --disable-fail-fast --disable-reuse-test --stage-list "DGX_B200-8_GPUs-PyTorch-4" |
|
PR_Github #60003 [ run ] triggered by Bot. Commit: |
|
PR_Github #60003 [ run ] completed with state
|
|
/bot run --disable-fail-fast --disable-reuse-test --stage-list "DGX_B200-8_GPUs-PyTorch-4" |
|
PR_Github #60022 [ run ] triggered by Bot. Commit: |
|
PR_Github #60022 [ run ] completed with state
|
|
/bot run --disable-fail-fast --disable-reuse-test --stage-list "DGX_B200-8_GPUs-PyTorch-4" |
|
PR_Github #60066 [ run ] triggered by Bot. Commit: |
|
PR_Github #60066 [ run ] completed with state
|
|
/bot run --disable-fail-fast --disable-reuse-test --stage-list "DGX_B200-8_GPUs-PyTorch-4" |
|
PR_Github #60081 [ run ] triggered by Bot. Commit: |
|
PR_Github #60081 [ run ] completed with state
|
|
/bot run --disable-fail-fast --disable-reuse-test --stage-list "DGX_B200-8_GPUs-PyTorch-4" |
|
PR_Github #60089 [ run ] triggered by Bot. Commit: |
|
PR_Github #60089 [ run ] completed with state |
|
/bot run --disable-fail-fast --stage-list "DGX_B200-8_GPUs-PyTorch-4" |
|
/bot run --only-qa-verify test accuracy/test_disaggregated_serving.py::TestNemotron3Super120B::test_ctx_dp2_gen_tp4 |
|
PR_Github #61143 [ run ] triggered by Bot. Commit: |
|
PR_Github #61143 [ run ] completed with state |
|
/bot run --disable-fail-fast |
|
PR_Github #61178 [ run ] triggered by Bot. Commit: |
|
PR_Github #61178 [ run ] completed with state
|
|
/bot run --disable-fail-fast |
|
PR_Github #61311 [ run ] triggered by Bot. Commit: |
|
PR_Github #61311 [ run ] completed with state
|
|
/bot run --disable-fail-fast |
|
PR_Github #61394 [ run ] triggered by Bot. Commit: |
|
PR_Github #61394 [ run ] completed with state
|
|
/bot run --disable-fail-fast |
|
PR_Github #61597 [ run ] triggered by Bot. Commit: |
|
PR_Github #61597 [ run ] completed with state
|
|
/bot run --disable-fail-fast |
|
PR_Github #61623 [ run ] triggered by Bot. Commit: |
|
PR_Github #61623 [ run ] completed with state
|
|
/bot run --disable-fail-fast |
|
PR_Github #61661 [ run ] triggered by Bot. Commit: |
|
PR_Github #61661 [ run ] completed with state |
…ated transfer (NVIDIA#16505) Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com> Co-authored-by: xinhe-nv <200704525+xinhe-nv@users.noreply.github.com>
…nnahz/dep-1083-port-flashinfer-stable-va-lifecycle-for-native-all-reduce * 'main' of https://github.com/NVIDIA/TensorRT-LLM: (54 commits) [NVIDIA#15673][fix] Enable CUDA core fast path for SM89/SM120/SM121 (NVIDIA#12705) [None][test] Adjust timeout cases in QA perf test (NVIDIA#16894) [https://nvbugs/6157892][fix] Mistral format refactor (NVIDIA#15123) [None][feat] Add kimi_k2/glm_5 grouped routing and fused router to bench_moe (NVIDIA#16830) [https://nvbugs/6501376][fix] Test-only fix — drop the `if hidden_size % 2 != 0: with pytest.raises(...)`… (NVIDIA#16844) [TRTLLM-13642][feat] Add perf sanity tests for Llama-3.1-8B and Gemma-3-1B and verify cache transceiver V2 support (NVIDIA#16355) [https://nvbugs/6433376][fix] Update the Dense test to mirror the MoE sibling — assert `bfloat16` under… (NVIDIA#16203) [None][fix] Resolve NVFP4 mixed-precision base layers for the DSpark draft (NVIDIA#16831) [https://nvbugs/6479324][test] Remove waiver for fixed qwen3_5_4b_fp8_stress disaggregated stress test (NVIDIA#16878) [https://nvbugs/6507109][infra] Split slow DGX B300 attention unit tests (NVIDIA#16838) [None][infra] Waive 21 failed cases for main in post-merge 2862 (NVIDIA#16882) [None][perf] prepare_inputs: avoid O(seq_len) get_tokens(0) marshalling on the host (NVIDIA#16791) [None][perf] Optimize Blackwell fused MHC half-MMA kernel (NVIDIA#16799) [None][infra] Auto-update test durations from OpenSearch (last 7 days) [None][perf] Skip DeepGEMM clean_logits in DSA indexer prefill on custom top-k path (NVIDIA#16789) [None][feat] Support DeepSeek-V4 in layer_wise_benchmarks (NVIDIA#16774) [https://nvbugs/6465993][fix] use attention cache dtype for disaggregated transfer (NVIDIA#16505) [https://nvbugs/6463822][fix] Fix LTX2 CUDA graph test leak issue (NVIDIA#16775) [https://nvbugs/5948435][chore] Unwaive DeepSeekV3Lite test_nvfp4_4gpus CUTLASS ep4 fp8kv on RTXPro6000D (NVIDIA#16621) [TRTLLM-14417][fix] Exclude ADP/cuda-graph dummy requests from speculative-decode acceptance stats (NVIDIA#16571) ... Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
…nnahz/dep-1082-shared-mnnvl-moe-lifecycle * 'main' of https://github.com/NVIDIA/TensorRT-LLM: (142 commits) [NVIDIA#15673][fix] Enable CUDA core fast path for SM89/SM120/SM121 (NVIDIA#12705) [None][test] Adjust timeout cases in QA perf test (NVIDIA#16894) [https://nvbugs/6157892][fix] Mistral format refactor (NVIDIA#15123) [None][feat] Add kimi_k2/glm_5 grouped routing and fused router to bench_moe (NVIDIA#16830) [https://nvbugs/6501376][fix] Test-only fix — drop the `if hidden_size % 2 != 0: with pytest.raises(...)`… (NVIDIA#16844) [TRTLLM-13642][feat] Add perf sanity tests for Llama-3.1-8B and Gemma-3-1B and verify cache transceiver V2 support (NVIDIA#16355) [https://nvbugs/6433376][fix] Update the Dense test to mirror the MoE sibling — assert `bfloat16` under… (NVIDIA#16203) [None][fix] Resolve NVFP4 mixed-precision base layers for the DSpark draft (NVIDIA#16831) [https://nvbugs/6479324][test] Remove waiver for fixed qwen3_5_4b_fp8_stress disaggregated stress test (NVIDIA#16878) [https://nvbugs/6507109][infra] Split slow DGX B300 attention unit tests (NVIDIA#16838) [None][infra] Waive 21 failed cases for main in post-merge 2862 (NVIDIA#16882) [None][perf] prepare_inputs: avoid O(seq_len) get_tokens(0) marshalling on the host (NVIDIA#16791) [None][perf] Optimize Blackwell fused MHC half-MMA kernel (NVIDIA#16799) [None][infra] Auto-update test durations from OpenSearch (last 7 days) [None][perf] Skip DeepGEMM clean_logits in DSA indexer prefill on custom top-k path (NVIDIA#16789) [None][feat] Support DeepSeek-V4 in layer_wise_benchmarks (NVIDIA#16774) [https://nvbugs/6465993][fix] use attention cache dtype for disaggregated transfer (NVIDIA#16505) [https://nvbugs/6463822][fix] Fix LTX2 CUDA graph test leak issue (NVIDIA#16775) [https://nvbugs/5948435][chore] Unwaive DeepSeekV3Lite test_nvfp4_4gpus CUTLASS ep4 fp8kv on RTXPro6000D (NVIDIA#16621) [TRTLLM-14417][fix] Exclude ADP/cuda-graph dummy requests from speculative-decode acceptance stats (NVIDIA#16571) ... Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
Description
Fix disaggregated KV-cache transport for hybrid recurrent/attention models and remove the three NVBug 6465993 Nemotron 3 Super waivers.
Root cause
CppMambaHybridCacheManagerkeeps separateHALFrecurrent-state andFP8attention KV pools.CacheTransBufferManagerselected its wire dtype and size through layer-indexedgetPrimaryPool(0), which resolved to the recurrent pool whileCacheFormattertransferred attention blocks. The resultingHALFtransport buffer versusFP8payload mismatch failed on the first request and surfaced as a timeout. Controls reproduced the failure with session reuse disabled and with the worker monitor bypassed, ruling both out.Regression history
Fix
BlockManagerpools, excluding block-scale and indexer pools.Verification
f0f316bd:test_ctx_dp2_gen_tp4— passed inDGX_B200-8_GPUs-PyTorch-1(#48500).test_auto_dtype[mtp_nextn=0-block_reuse=False-use_py_transceiver=False]— passed inDGX_B200-8_GPUs-PyTorch-2(#48500).test_auto_dtype[mtp_nextn=3-block_reuse=True-use_py_transceiver=False]— passed inDGX_B200-8_GPUs-PyTorch-4(#48486); it also passed independently on the preceding fixed head in #48476.f0f316bd.94ce31a9. Subsequent full runs did not reach the B200 multi-GPU coverage: #48566 was blocked by the upstream DSpark regression repaired in #16579, while #48717 and #48756 were blocked by unrelated CI/baseline instability before multi-GPU dispatch. No reported failure implicated this PR's code.Dev Engineer Review
CacheTransBufferManagerto derive the wire/transferDataTypefrom actual KV-cache pool metadata (preferring attention-pool dtype when present, enforcing consistent attention-pool dtype, and using indexer-K pool dtype whentransferIndexerKCacheis enabled).CacheFormatter::unformatto allocate its temporary receive buffer usingmCacheTransBufferManager->getDataType()rather than deriving dtype from the source cache’s pool-0.CacheTransBufferManager::computeTransferBufferSize, corrected non-indexer-K sizing to compute the byte-per-token-per-layer term per layer from that layer’s pool primary dimensions (instead of using a single pool-0-derived value). Indexer-K sizing behavior remains indexer-K dimension-based.CacheTransBufferManagerthrowstensorrt_llm::common::TllmExceptionwhen eligible attention-pool dtypes are mixed.CacheTransBufferManager::getDataType()accessor to expose the selected transfer dtype for downstream buffer allocations.QA Engineer Review
TEST_F(KVCacheManagerTest, HybridDisaggUsesAttentionPoolDtype)CacheTransBufferManager::getDataType(),getSendBuffer(...), and consistent buffer index handling viafreeBufferIndexForSend).TEST_F(KVCacheManagerTest, VswaDisaggDtypeMismatchTriggersGuard)tests/integration/test_lists/:tests/integration/test_lists/waives.txtaccuracy/test_disaggregated_serving.py::TestNemotron3Super120B::test_auto_dtype[mtp_nextn=0-block_reuse=False-use_py_transceiver=True](nvbugs/6478726)test_auto_dtypevariants forTestNemotron3Super120B(including theuse_py_transceiver=Falseandmtp_nextn=3-block_reuse=True-use_py_transceiver=Falsecases) and removed the separateaccuracy/test_disaggregated_serving.py::TestNemotron3Super120B::test_ctx_dp2_gen_tp4waive entry (all tied to nvbugs/6465993).