Skip to content

Repository files navigation

🧭 AdaSpatial-MLLM

Training-Free Adaptive Reasoning for 3D Spatial MLLMs

面向 3D 空间理解的自适应推理 MLLM

Training Free RTX 5090 Python 3.10+ PyTorch 2.7+ Tests Artifacts License

A reproducible research engineering project built on the released 3DThinker Stage-1 checkpoint. It studies when latent 3D reasoning can stop early—and when difficult samples should receive extra inference—without retraining model weights.

Results · Protocol · Interactive demo

3DThinker-R training-free adaptive latent reasoning project overview

Core adaptive early-exit pipeline and confirmed MindCube-Tiny holdout results.

✨ What is new

  • Adaptive latent early exit based on consecutive hidden-state convergence.
  • Task-aware routing for Among, Rotation, and Around spatial question families.
  • Adaptive+ difficult-sample cascade using fixed-4, fixed-8, and adaptive majority voting.
  • Answer-confidence diagnostics and an optional left/right mirror-consistency check.
  • Auditable inference traces: answer, latent trajectory, exit reason, latency, and peak VRAM.
  • Real GPU demo for 2–4 uploaded views, tested on one NVIDIA RTX 5090.

No fine-tuning is required. The released Stage-1 checkpoint remains frozen.

🧠 Method

AdaSpatial-MLLM Adaptive+ architecture

The default path runs the frozen 3D MLLM and exits after two consecutive latent-state cosine distances fall below ε=0.04. A task router can preserve the full budget where early exit is less reliable. For a difficult Among sample that reaches the maximum budget without convergence, Adaptive+ adds fixed-4 and fixed-8 predictions and returns a deterministic majority vote.

The optional symmetry branch horizontally flips directional inputs and maps left/right answers back before checking consistency. It is exposed as a diagnostic and is not included in the main 68.53% exploratory result.

📊 Results

All accuracy numbers use the released 3B Stage-1 checkpoint, greedy decoding, seed 42, batch size 1, and a frozen 950-example MindCube-Tiny holdout. See the full evaluation protocol and machine-readable summary files.

Confirmatory evaluation

Policy Accuracy Mean latent steps Early exit
Fixed-4 66.00% 4.00 —
Fixed-8 63.16% 8.00 —
Fixed-12 67.16% 11.97 —
Adaptive 66.95% 9.69 57.8%

Adaptive preserved held-out accuracy relative to fixed-12 (−0.21 percentage points; exact McNemar p=0.935) while reducing mean latent steps by 19.0%.

Accuracy versus latent steps

Exploratory Adaptive+

Policy Accuracy Mean inference latent steps Escalated
Adaptive 66.95% 9.69 —
Task-aware routing 68.11% 10.54 —
Adaptive+ cascade 68.53% 12.99 194 / 950

Adaptive+ corrected 15 net errors relative to Adaptive (+1.58 points), but the paired difference is not statistically significant (p=0.115). This policy was designed after inspecting the main prediction matrix, so it is reported honestly as post-hoc exploratory evidence, not a new state-of-the-art claim. A fresh frozen test set is required for confirmation.

Task-level accuracy chart Single-process latency chart

Clean single-process timing measured 6.104 s for Adaptive and 6.306 s for fixed-12 on 50 samples (12.64 GiB peak VRAM). Latent-step savings do not translate one-for-one into wall-clock savings because answer decoding dominates runtime.

🖥️ Live GPU demo

The web interface offers two clearly separated modes:

  • Replay: dependency-free playback of checked-in, real evaluation traces.
  • Live GPU: upload 2–4 views and compare fixed-12 with Adaptive+ on the actual checkpoint.

Live RTX 5090 Adaptive+ demo correcting a difficult spatial sample

Uploaded multi-view input and spatial question Adaptive latent-state trace on RTX 5090

The screenshot above is a real RTX 5090 run. On this difficult camera-motion sample, the normal 12-step path selected the wrong option while the conditional fixed-4/fixed-8 vote corrected the final answer. This is a representative success case—not aggregate evidence; the full 950-sample results remain the primary evaluation.

Start the interface

python -m http.server 7860 --directory demo

Open http://localhost:7860. Replay mode works immediately.

Enable real inference

Use the environment containing the patched Stage-1 dependencies and downloaded checkpoint:

pip install -e ".[demo]"
export THINKER_MODEL_PATH=/path/to/3DThinker-Mindcube
python -m uvicorn demo.live_server:app --host 127.0.0.1 --port 8001

For a remote GPU, keep the API private and forward it over SSH:

ssh -N -L 8001:127.0.0.1:8001 -p <port> <user>@<gpu-server>

The first request loads the checkpoint; later requests reuse the resident model. GPU requests are serialized to prevent concurrent jobs from competing for VRAM. See demo/README.md for endpoint details and upload limits.

🧪 Reproduce the research artifacts

Install the lightweight analysis package:

python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev,report]"
pytest -q

Validate every checked-in result and regenerate the summaries:

python scripts/validate_artifacts.py
python scripts/summarize_results.py
python scripts/analyze_research_results.py
python scripts/evaluate_optimized_router.py

Full GPU inference requires the original Stage-1 environment, MindCube images, and released model checkpoint. Example commands and frozen settings are documented in docs/EXPERIMENT_PROTOCOL.md.

📦 Data and artifact policy

Artifact Included Location
Aggregate CSV/JSON results and statistical tests ✅ results/summary
Publication-ready charts ✅ assets
Curated demo traces and sample views ✅ demo/data, demo/assets
Architecture diagram ✅ assets/adaptiveplus_architecture.svg
Raw benchmark images and model weights ❌ Download from the original providers
Per-sample raw model generations ❌ Excluded from the public Git history

Demo sample images originate from MindCube and remain subject to the original dataset terms. Raw MindCube data and model weights are intentionally not redistributed to keep the repository small and licensing boundaries clear.

🗂️ Repository map

src/thinker_r/          adaptive policy, routing, interventions, metrics
demo/                   replay UI and FastAPI live-GPU backend
scripts/                inference, analysis, timing, and artifact validation
tests_r/                focused tests for the added research code
results/summary/        compact machine-readable evidence
reports/                human-readable results and audit report
assets/                 architecture and result figures
3dthinker/              attributed upstream 3DThinker implementation

⚠️ Scope and limitations

  • Evaluated on one released 3B Stage-1 checkpoint and MindCube-Tiny.
  • Adaptive+ is exploratory; its apparent accuracy improvement needs fresh confirmatory data.
  • Answer-token confidence is relative over A–D forms, not a calibrated correctness probability.
  • Fixed-duration latent interventions did not show a uniform accuracy collapse; multi-seed replication is still needed for random controls.
  • This repository contributes inference, evaluation, and demo engineering—not a newly trained foundation model.

🙏 Upstream attribution

AdaSpatial-MLLM is an independent research extension of zhangquanchen/3DThinker. The original model, training code, and scientific contribution belong to Zhangquan Chen, Manyuan Zhang, Xinlei Yu, Xufang Luo, Mingze Sun, Zihao Pan, Yan Feng, Peng Pei, Xunliang Cai, and Ruqi Huang.

If you use the base model or code, cite the original 3DThinker work and follow its model/data licenses. This derivative repository retains the upstream Apache-2.0 license.

📄 License

Licensed under the Apache License 2.0.

About

面向3D空间理解的自适应推理MLLM

Topics

Resources

Code of conduct

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages