Skip to content

[NVBUG-6327718][test] Unwaive test_disaggregated_videomme[nemotron_na… - #17031

Draft
aswinvisva wants to merge 1 commit into
NVIDIA:mainfrom
aswinvisva:avisva/unwaive-nvbug-6327718
Draft

[NVBUG-6327718][test] Unwaive test_disaggregated_videomme[nemotron_na…#17031
aswinvisva wants to merge 1 commit into
NVIDIA:mainfrom
aswinvisva:avisva/unwaive-nvbug-6327718

Conversation

@aswinvisva

@aswinvisva aswinvisva commented Jul 29, 2026

Copy link
Copy Markdown
Collaborator

…no_v3_omni_fp8] on B200 and H20

The root cause of the pre-fix failure — TypeError: isinstance() arg 2 must be a type, a tuple of types, or a union at tensorrt_llm/executor/proxy.py when the test infrastructure's session-reuse cache monkey-patches MpiPoolSession in the proxy module namespace with a factory function — is already fixed on main. The proxy code was refactored to identify pool-backed sessions by excluding (MpiCommSession, RemoteMpiCommSessionClient) instead of isinstance(x, MpiPoolSession).

Remove the two waivers linked to NVBug 6327718 so CI can verify the fix on both platforms.

Dev Engineer Review

  • Removed the two nvbug: 6327718 waivers for the targeted test on B200 and H20.
  • The change is correctly scoped to tests/integration/test_lists/waives.txt with no code or API changes.
  • Waiver formatting and platform coverage should remain consistent with the surrounding entries.

QA Engineer Review

  • Test-list-only change; no test-db/ or qa/ files were modified.
  • Removed CI waiver entries for TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_fp8] on full:B200 and full:H20.
  • Verdict: needs follow-up pending CBTS coverage data.

Description

Test Coverage

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

@aswinvisva
aswinvisva requested review from a team as code owners July 29, 2026 22:47
@aswinvisva
aswinvisva marked this pull request as draft July 29, 2026 22:48
@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_B200-PyTorch-4"

@coderabbitai

coderabbitai Bot commented Jul 29, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 379b5893-92c2-4acb-9e35-80cf7e46358d

📥 Commits

Reviewing files that changed from the base of the PR and between c45ad83 and 08015b1.

📒 Files selected for processing (1)
  • tests/integration/test_lists/waives.txt
💤 Files with no reviewable changes (1)
  • tests/integration/test_lists/waives.txt

Walkthrough

Two waiver entries are removed from the integration test waiver list, allowing the disaggregated VideoMME test to run on B200 and H20.

Changes

Test waiver cleanup

Layer / File(s) Summary
Remove VideoMME waivers
tests/integration/test_lists/waives.txt
Removes SKIP entries for the targeted test on full:B200 and full:H20.

Estimated code review effort: 1 (Trivial) | ~2 minutes

Possibly related PRs

Suggested reviewers: qijune, brnguyen2

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description explains the change, but the required Test Coverage and PR Checklist sections are left empty. Add a brief Test Coverage note and complete the PR Checklist items, or explain why they do not apply.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title is specific and matches the main change: removing a waiver for the affected test.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #62622 [ run ] triggered by Bot. Commit: 08015b1 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #62622 [ run ] completed with state SUCCESS. Commit: 08015b1
/LLM/main/L0_MergeRequest_PR pipeline #50762 (Partly Tested) completed with status: 'SUCCESS'

CI Report

Link to invocation

@BowenFu

BowenFu commented Aug 1, 2026

Copy link
Copy Markdown

The two lines removed are full:-scoped, i.e. QA/post-merge lists. The green /bot run --stage-list "DGX_B200-PyTorch-4" here executes l0_b200.yml, which lists only the qwen3vl_2b_instruct and nemotron_nano_v3_omni_nvfp4 variants — not the fp8 one being unwaived — so this PR's CI is zero coverage for the change.

Two other gaps against the description: nvbugs/6327718 is still P0 Dev - Open - To fix, and its recorded signature is Fatal Python error: Aborted in executor/ipc.py / inputs/multimodal_data.py, not the isinstance() TypeError; the repair-bot cannot-reproduce comment is for the nvfp4 variant. And the fix you cite (e15883702b, #16444) merged 2026-07-16, while these two waivers were added on 07-19/20 by #16593 / #16602 — after it.

Could you attach a passing run of the fp8 variant on B200 and H20 (host + commit + pytest command), the way repair-bot did for nvfp4? For what it's worth the fp8 variant already runs unwaived pre-merge via l0_h100.yml:166, so the removal is probably right — it just isn't demonstrated on the two platforms being unwaived.

@brnguyen2 brnguyen2 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks like a no-op. [tests/integration/test_lists/waives.txt:16](https://github.com/NVIDIA/TensorRT-LLM/pull/17031/files#diff-621bd2af82a3b97c7a5948368d36c14582ffdf73361deb7f42d2ff395b22167eR16) still waives the same node ID unconditionally (accuracy/test_epd_disagg_multimodal.py::TestVideoMMEEPD::test_disaggregated_videomme[nemotron_nano_v3_omni_fp8], nvbugs/6478692), so the test stays skipped on every platform and CI can't confirm the fix. Also, the fp8 param only appears in test-db/l0_h100.yml:164 and qa/llm_function_core.txt:86l0_b200.yml lists only the qwen3vl/nvfp4 variants — so the B200 entry was already vestigial. Worth confirming the 6327718 signature matches what you believe is fixed, and updating that bug alongside the removal.

…no_v3_omni_fp8] and stress-repeat 20x

Removes two waivers linked to NVBug 6327718 for the Nemotron Nano V3 Omni fp8
E/PD-disagg VideoMME test on B200 and H20. The recorded signature on 6327718
is a stochastic `Fatal Python error: Aborted` (SIGABRT) that appeared in about
2.5%% of pre-merge runs from 2026-06-09 through 2026-07-19 under the 128-way
ThreadPoolExecutor in the videomme driver. This test is currently believed to
be fixed on main; this PR is a stress-test to demonstrate that.

Changes:
  * tests/integration/test_lists/waives.txt: drop two `full:B200/` and
    `full:H20/` skip entries for 6327718.
  * tests/integration/defs/accuracy/test_epd_disagg_multimodal.py: add a
    `_repeat` parametrize dimension so each variant runs 20 iterations per
    stage. Because the SIGABRT is stochastic, a single passing run does not
    beat the ~2.5%% observed rate; 20 back-to-back reps give ~40%% detection
    probability if the bug were still latent.
  * tests/integration/test_lists/test-db/l0_h100.yml: enumerate 20 rep IDs
    for the fp8 variant (the only one wired into H100 pre-merge).
  * tests/integration/test_lists/test-db/l0_b200.yml: enumerate 20 rep IDs
    for the nvfp4 variant; keep qwen3vl_2b_instruct at a single rep so it
    still gets collected on B200.
  * tests/integration/test_lists/qa/llm_function_core.txt: update all three
    variants to the new `-rep00` naming for the QA collection.

This is a throwaway PR intended for stress verification, not merge.

Trigger with:
  /bot run --stage-list "DGX_H100-PyTorch-1, DGX_B200-PyTorch-4" --disable-fail-fast

Related: NVBug 6478692 (a separate, previously-global waiver on the same
node ID) has since been removed on main; no reference remains in waives.txt.

Signed-off-by: Aswin Visva <31215515+aswinvisva@users.noreply.github.com>
@aswinvisva
aswinvisva force-pushed the avisva/unwaive-nvbug-6327718 branch from 08015b1 to eefdfc4 Compare August 10, 2026 16:25
@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_H100-PyTorch-1, DGX_B200-PyTorch-4" --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65073 [ run ] triggered by Bot. Commit: eefdfc4 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65073 [ run ] completed with state FAILURE. Commit: eefdfc4
/LLM/main/L0_MergeRequest_PR pipeline #52878 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@aswinvisva

Copy link
Copy Markdown
Collaborator Author

/bot run --stage-list "DGX_H100-PyTorch-1, DGX_B200-PyTorch-4" --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65111 [ run ] triggered by Bot. Commit: eefdfc4 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #65111 [ run ] completed with state SUCCESS. Commit: eefdfc4
/LLM/main/L0_MergeRequest_PR pipeline #52909 (Partly Tested) completed with status: 'SUCCESS'

CI Report

Link to invocation

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants