Yuehao Huang1,2,3,β , Yunzi Wu1,2,4,β , Xiaotao Zhang1,2,5,β , Xinhai Li1,2,β‘, Jiankun Dong1,2, Jiajun Lv3, Chi Zhang1,2, Chenjia Bai1,2,*, Yong Liu3,*, Xuelong Li1,2,*
1 Institute of Artificial Intelligence, China Telecom Β 2 Gamma Robotics Β 3 Zhejiang University Β 4 Tongji University Β 5 Shanghai Jiao Tong University
β Equal Contributions Β β‘ Project Leader Β * Corresponding Authors
WNM-3D-RL is the official Stage-III reinforcement-learning companion to WNM-3D. Built on VERL-Omni, it provides reward-guided post-training for WNM-3D's joint future-view and navigation-action generator.
This repository contains the paper-aligned Counterfactual DanceGRPO recipe, WNM-3D rollout and actor integrations, navigation rewards, GN0 data conversion, and checkpoint export tools. Stage-I A* supervision, Stage-II DAgger-SFT, released checkpoints, and online inference are maintained in the main WNM-3D repository.
Important
This repository does not include WNM-3D weights, training data, simulator assets, or the external WNM-3D model implementation. Prepare Stage-I and Stage-II with WNM-3D before starting Stage-III here.
WNM-3D follows a progressive three-stage training curriculum:
| Stage | Method | Training signal | Repository |
|---|---|---|---|
| I | Offline A* SFT | Expert future views and actions | WNM-3D |
| II | Closed-loop DAgger-SFT | Expert corrections at policy-visited states | WNM-3D + GN0 |
| III | Counterfactual DanceGRPO | Visual, navigation, STOP, yaw, and collision rewards | This repository |
Stage-III starts from a complete Stage-II WNM-3D checkpoint. It samples counterfactual future-view/action trajectories from fixed DAgger contexts and optimizes them with decomposed GRPO advantages. The simulator is not stepped during optimization; GN0 records provide the reference trajectory, occupancy, goal, and event metadata used by the reward functions.
- Joint world-action policy optimization. The trainable video/action DiT is updated while the WNM-3D geometry encoders and other pretrained components remain frozen.
- Counterfactual stratified rollout. Each prompt uses four diffusion-noise strata and four counterfactual branches per stratum, with shared initial noise and deterministic request-level RNG.
- Layer-conditioned credit assignment. Visual quality, all four navigation chunks, and all four STOP chunks receive separately normalized GRPO advantages. Collision penalties and recovery shaping remain within the navigation reward.
- Deployment-aligned reward design. The official recipe combines distance-scaled premature-STOP penalties; path-, rate-, and reference-heading yaw guards; hard-collision penalties; soft clearance-risk shaping; and bounded collision-recovery bonuses.
- Fail-closed release contract. Dataset provenance, checkpoint schema, action normalization, topology divisibility, and algorithm constants are validated before Ray starts.
The sole supported recipe is wnm-3d-stage3-navigation-grpo-r1:
bash recipes/wnm_3d/stage3/launch.sh[2026-08-25]We released the official Stage-III Counterfactual DanceGRPO training code for WNM-3D.[2026-08-12]The official WNM-3D project website was launched.[2026-08-07]The WNM-3D paper became available on arXiv.
- π Overview
- π News
- π Installation
- π€ Workspace and Artifacts
- ποΈ Prepare Stage-III Data
- βοΈ Configure the Recipe
- π Launch Training
- π¦ Export a Checkpoint
- π Repository Layout
- π§ͺ Development
- π Citation
- π Acknowledgements
- π License
WNM-3D-RL requires Linux, NVIDIA GPUs, Python 3.11 or 3.12, a CUDA-compatible PyTorch stack, the pinned VERL revision, and the external WNM-3D checkout. The reference CUDA image and CI use Python 3.12.
Clone WNM-3D-RL next to WNM-3D:
git clone https://github.com/TeleHuman/WNM-3D-RL.git
cd WNM-3D-RLInstall the rollout and training dependencies in this order inside the target environment:
conda create -n wnm3d python=3.12 -y
conda activate wnm3d
export UV_PROJECT_ENVIRONMENT="$CONDA_PREFIX"
python -m pip install --upgrade uv
uv pip install -e ".[gpu]" --torch-backend=auto
uv pip install \
"vllm-omni @ git+https://github.com/vllm-project/vllm-omni.git@$(cat .github/vllm_omni_pin.txt)"
uv pip install -e ".[train,wam,dev]"
MAX_JOBS=8 uv pip install --no-build-isolation -e ".[attention]"uv pip install already discovers an activated Conda environment through
CONDA_PREFIX; setting UV_PROJECT_ENVIRONMENT additionally keeps uv project
commands such as uv sync and uv run in that same environment.
The verl-omni wheel contains only the verl_omni Python library. The
official Stage-III recipe, cluster launch scripts, and runtime contracts are
source-tree assets; run complete training from a Git checkout rather than from
an installed wheel alone.
WNM-3D is consumed from WNM3D_SOURCE_ROOT through PYTHONPATH. Do not install
the WNM-3D project dependencies into this environment: its standalone SFT and
inference environment has different Torch, Transformers, and Diffusers pins.
Verify the core imports in the same environment that will launch Ray:
python -c 'import ray, torch, pyarrow, safetensors, verl, verl_omni'Alternatively, build the reference CUDA/Ray/RDMA image:
docker build -f docker/Dockerfile.cuda --target dev -t wnm-3d-rl:cu130-dev .The image intentionally excludes the external WNM-3D source, model weights, datasets, and outputs. Mount those resources at runtime.
For end-to-end data preparation, the recommended workspace keeps the model, RL, and benchmark repositories as siblings:
workspace/
βββ WNM-3D/ Stage-I/II training, model implementation, and inference
βββ WNM-3D-RL/ Stage-III reinforcement learning
βββ GN0/ DAgger collection and GN-Bench evaluation
GN0 is required only to collect or regenerate the Stage-III DAgger/reference records. If the converted Parquet files, referenced source videos, and occupancy assets are already available, Stage-III training itself does not require the GN0 checkout or the GN-Bench environment.
Stage-III training requires:
- the external WNM-3D source checkout;
- a complete Stage-II WNM-3D checkpoint for policy initialization;
- the checkpoint used to define the data and action-normalization contract;
- UMT5 tokenizer and text-encoder weights;
- CLIP image-encoder weights;
- the Wan2.2 VAE;
- the VGGT-Ξ© checkpoint;
- Stage-III GN0 DAgger/reference records; and
- the corresponding InteriorGS occupancy assets.
Follow the main WNM-3D training guide to train Stage-I, collect DAgger records, and train Stage-II.
The initialization and data-contract checkpoints may contain different learned weights, but their model configuration, transform metadata, and action normalization must match exactly. The launcher verifies this contract.
Collect a dedicated Stage-III DAgger/reference root with the Stage-II policy as described by WNM-3D. Each converted sample contains 66 RGB frames: 33 history frames followed by 33 target frames. The records also retain the reference trajectory, goal, occupancy scene, action-normalization statistics, and event metadata required by the rewards.
The converter stores references to the source MP4 files rather than copying them. All training nodes must therefore resolve the same data paths.
export GN0_COLLECTION_ROOT=/path/to/gn0-stage3-collection
export OCCUPANCY_ROOT=/path/to/interiorgs-scenes
export DATA_CONTRACT_CKPT=/path/to/wnm-3d-stage2
export CONVERTED_ROOT=/path/to/converted-stage3-data
python tools/data/convert_interiorgs_dagger_to_rlhf.py \
--input-root "$GN0_COLLECTION_ROOT" \
--output-dir "$CONVERTED_ROOT" \
--checkpoint "$DATA_CONTRACT_CKPT" \
--occupancy-root "$OCCUPANCY_ROOT" \
--shard-size 2048 \
--val-fraction 0 \
--split-seed 42 \
--workers 32 \
--worker-prefetch 4Inspect manifest.json and require counts.skipped == 0. The manifest records
the model/config hashes and the complete action_normalization.json contract
inherited from the selected checkpoint. A CLI scale assertion is rejected if
it differs from the checkpoint; the converter never overrides the checkpoint
normalization.
Choose a row count compatible with the prompt batch for the target topology:
export TRAIN_ROOT=/path/to/stage3-training-data
export TRAIN_ROW_COUNT=<TRAIN_ROW_COUNT>
python tools/data/sample_stage3_train.py \
--source-manifest "$CONVERTED_ROOT/manifest.json" \
--output-dir "$TRAIN_ROOT" \
--train-size "$TRAIN_ROW_COUNT" \
--shard-size 2048 \
--seed 42Sampling is deterministic and keeps at most one record for each source episode, preventing query variants of the same trajectory from crossing the training contract.
export EVENT_VAL_ROOT=/path/to/stage3-event-validation
export SAMPLES_PER_EVENT=<SAMPLES_PER_EVENT>
python tools/data/build_vggt_stage3_event_val.py \
--input-root "$GN0_COLLECTION_ROOT" \
--output-dir "$EVENT_VAL_ROOT" \
--checkpoint "$DATA_CONTRACT_CKPT" \
--exclude-parquet-dir "$TRAIN_ROOT" \
--occupancy-root "$OCCUPANCY_ROOT" \
--per-event "$SAMPLES_PER_EVENT" \
--workers 32The held-out set balances collision precursors, premature-STOP risk, required STOP, and near-goal continuation while excluding training sources.
Copy the machine-local path template:
cp recipes/wnm_3d/stage3/paths.env.example \
recipes/wnm_3d/stage3/paths.envEdit paths.env and provide:
| Variable | Purpose |
|---|---|
WNM3D_SOURCE_ROOT |
WNM-3D source checkout |
WNM3D_INITIAL_CHECKPOINT |
Complete Stage-II initialization checkpoint |
WAM_DATA_CONTRACT_CHECKPOINT |
Checkpoint used during data conversion |
WNM_TOKENIZER_SOURCE |
UMT5 tokenizer directory |
WNM_TEXT_ENCODER_SOURCE |
Frozen UMT5 encoder weights |
WNM_IMAGE_ENCODER_SOURCE |
Frozen CLIP encoder weights |
WNM_VAE_SOURCE |
Wan2.2 VAE weights |
WNM_VGGT_SOURCE |
VGGT-Ξ© checkpoint |
WAM_DATA_ROOT |
Sampled Stage-III training dataset |
WAM_EVENT_VAL_FILE |
Event-balanced validation Parquet |
WAM_EXPECTED_TRAIN_SIZE |
Exact training row count from the manifest |
WAM_VAL_MAX_SAMPLES |
Exact validation row count to evaluate |
WAM_OUTPUT_DIR |
Durable output directory |
paths.env is ignored by Git. Do not commit machine paths, credentials,
private weights, datasets, logs, videos, or TensorBoard events.
Validate the complete data/model/topology contract before allocating the training job:
bash recipes/wnm_3d/stage3/validate.shA valid setup prints STAGE3_CONFIG_CONTRACT_OK.
The default topology is one node with eight GPUs:
bash recipes/wnm_3d/stage3/launch.shRun the same launcher on every homogeneous node with a common head address and a distinct zero-based rank.
Node 0:
NNODES=<NODE_COUNT> \
NODE_RANK=0 \
MASTER_ADDR=<HEAD_IPV4> \
bash recipes/wnm_3d/stage3/launch.shEach worker:
NNODES=<NODE_COUNT> \
NODE_RANK=<RANK> \
MASTER_ADDR=<HEAD_IPV4> \
bash recipes/wnm_3d/stage3/launch.shThe resource resolver preserves 16 rollouts per prompt and derives prompt, PPO, and micro-batches that satisfy exact FSDP divisibility. Positional Hydra overrides are intentionally rejected so they cannot bypass the verified algorithm contract.
The reference cluster uses bond0 for Ray/Gloo/NCCL bootstrap and eight
mlx5_* devices for NCCL IB/RoCE traffic. Override these values in paths.env
when the cluster topology differs. The launcher refuses to terminate an
existing Ray runtime unless WAM_REPLACE_EXISTING_RAY=true is explicitly set.
Training logs, TensorBoard events, resolved configuration snapshots, and
automatic checkpoints are written below WAM_OUTPUT_DIR:
tensorboard --logdir "$WAM_OUTPUT_DIR" --bind_allVERL actor shards are not directly loadable by the WNM-3D inference server.
The exporter merges the trained joint video/action DiT into the complete
Stage-II WNM-3D checkpoint while preserving frozen encoders, the VAE,
configuration, and the complete action_normalization.json contract.
Run a dry-run first:
export BASE_WNM=/path/to/wnm-3d-stage2
export ACTOR_CKPT=/path/to/stage3-output/global_step_N/actor
export EXPORTED_WNM=/path/to/exported-wnm-3d-stage3
python tools/checkpoint/export_wnm_checkpoint.py \
--base-vln "$BASE_WNM" \
--actor-checkpoint "$ACTOR_CKPT" \
--output-dir "$EXPORTED_WNM" \
--dry-runAfter checking tensor names, counts, and shapes, repeat without --dry-run:
python tools/checkpoint/export_wnm_checkpoint.py \
--base-vln "$BASE_WNM" \
--actor-checkpoint "$ACTOR_CKPT" \
--output-dir "$EXPORTED_WNM"The output directory must not already exist. Native FSDP export may require a
large local temporary disk; set TMPDIR=/path/to/local-tmp when needed.
For a one-off snapshot of a live FSDP run, use
tools/checkpoint/hot_save_verl_fsdp.py. Run its read-only preflight first and
add --execute only after verifying the Ray namespace, trainer PID, worker
count, and destination capacity.
recipes/wnm_3d/stage3/ official Stage-III recipe and launch contract
tools/data/ GN0 conversion and deterministic sampling
tools/checkpoint/ live FSDP snapshot and full-WNM export
verl_omni/pipelines/wnm_3d/ WNM-3D actor and rollout integration
verl_omni/pipelines/wnm_2d/ WNM-2D baseline integration
verl_omni/pipelines/wnm_shared/
shared WNM rollout mechanics
verl_omni/trainer/diffusion/ GRPO/PPO, replay, metrics, and layer credit
verl_omni/utils/reward_score/
navigation, STOP, collision, vision, and flow rewards
tests/ CPU and standalone regression tests
docker/Dockerfile.cuda CUDA/Ray/RDMA training image
Install the development dependencies and run the retained checks:
pre-commit install
ruff check .
ruff format --check .
find recipes/wnm_3d/stage3 -type f -name '*.sh' -print0 \
| xargs -0 -n1 bash -n
pytest -q $(find tests -type f -name '*_on_cpu.py' -print | sort)Run scripts/generate_trainer_config.sh after changing Hydra configuration.
If WNM-3D or this Stage-III implementation is useful for your research, please cite the WNM-3D paper:
@article{huang2026wnm_3d,
title={WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN},
author={Huang, Yuehao and Wu, Yunzi and Zhang, Xiaotao and Li, Xinhai and Dong, Jiankun and Lv, Jiajun and Zhang, Chi and Bai, Chenjia and Liu, Yong and Li, Xuelong},
journal={arXiv preprint arXiv:2608.07267},
year={2026}
}WNM-3D-RL builds on several open-source projects. We thank their authors and contributors:
- VERL and VERL-Omni provide the distributed reinforcement-learning foundation.
- DreamZero introduced the world-action modeling foundation used by WNM-3D.
- Wan2.2 provides the video diffusion backbone and pretrained video components.
- VGGT-Ξ© provides the frozen geometry encoder used for 3D scene conditioning.
- DINOv3 provides vision components incorporated through the WNM-3D geometry encoder.
- GN0/GN-Bench provides DAgger collection and closed-loop navigation evaluation.
Unless otherwise stated, source code in this repository is released under the Apache License 2.0. Third-party code, model weights, datasets, and generated artifacts remain subject to their respective licenses and terms.
The current complete WNM-3D Stage-III runtime depends on VGGT-Ξ© and is limited to the noncommercial research uses permitted by its license. Review THIRD_PARTY_NOTICES.md before redistributing code or model artifacts.
