面向 3D 空间理解的自适应推理 MLLM
A reproducible research engineering project built on the released 3DThinker Stage-1 checkpoint. It studies when latent 3D reasoning can stop early—and when difficult samples should receive extra inference—without retraining model weights.
Core adaptive early-exit pipeline and confirmed MindCube-Tiny holdout results.
- Adaptive latent early exit based on consecutive hidden-state convergence.
- Task-aware routing for Among, Rotation, and Around spatial question families.
- Adaptive+ difficult-sample cascade using fixed-4, fixed-8, and adaptive majority voting.
- Answer-confidence diagnostics and an optional left/right mirror-consistency check.
- Auditable inference traces: answer, latent trajectory, exit reason, latency, and peak VRAM.
- Real GPU demo for 2–4 uploaded views, tested on one NVIDIA RTX 5090.
No fine-tuning is required. The released Stage-1 checkpoint remains frozen.
The default path runs the frozen 3D MLLM and exits after two consecutive latent-state cosine
distances fall below ε=0.04. A task router can preserve the full budget where early exit is less
reliable. For a difficult Among sample that reaches the maximum budget without convergence,
Adaptive+ adds fixed-4 and fixed-8 predictions and returns a deterministic majority vote.
The optional symmetry branch horizontally flips directional inputs and maps left/right answers back before checking consistency. It is exposed as a diagnostic and is not included in the main 68.53% exploratory result.
All accuracy numbers use the released 3B Stage-1 checkpoint, greedy decoding, seed 42, batch size 1, and a frozen 950-example MindCube-Tiny holdout. See the full evaluation protocol and machine-readable summary files.
| Policy | Accuracy | Mean latent steps | Early exit |
|---|---|---|---|
| Fixed-4 | 66.00% | 4.00 | — |
| Fixed-8 | 63.16% | 8.00 | — |
| Fixed-12 | 67.16% | 11.97 | — |
| Adaptive | 66.95% | 9.69 | 57.8% |
Adaptive preserved held-out accuracy relative to fixed-12 (−0.21 percentage points; exact
McNemar p=0.935) while reducing mean latent steps by 19.0%.
| Policy | Accuracy | Mean inference latent steps | Escalated |
|---|---|---|---|
| Adaptive | 66.95% | 9.69 | — |
| Task-aware routing | 68.11% | 10.54 | — |
| Adaptive+ cascade | 68.53% | 12.99 | 194 / 950 |
Adaptive+ corrected 15 net errors relative to Adaptive (+1.58 points), but the paired difference
is not statistically significant (p=0.115). This policy was designed after inspecting the main
prediction matrix, so it is reported honestly as post-hoc exploratory evidence, not a new
state-of-the-art claim. A fresh frozen test set is required for confirmation.
Clean single-process timing measured 6.104 s for Adaptive and 6.306 s for fixed-12 on 50 samples (12.64 GiB peak VRAM). Latent-step savings do not translate one-for-one into wall-clock savings because answer decoding dominates runtime.
The web interface offers two clearly separated modes:
- Replay: dependency-free playback of checked-in, real evaluation traces.
- Live GPU: upload 2–4 views and compare fixed-12 with Adaptive+ on the actual checkpoint.
The screenshot above is a real RTX 5090 run. On this difficult camera-motion sample, the normal 12-step path selected the wrong option while the conditional fixed-4/fixed-8 vote corrected the final answer. This is a representative success case—not aggregate evidence; the full 950-sample results remain the primary evaluation.
python -m http.server 7860 --directory demoOpen http://localhost:7860. Replay mode works immediately.
Use the environment containing the patched Stage-1 dependencies and downloaded checkpoint:
pip install -e ".[demo]"
export THINKER_MODEL_PATH=/path/to/3DThinker-Mindcube
python -m uvicorn demo.live_server:app --host 127.0.0.1 --port 8001For a remote GPU, keep the API private and forward it over SSH:
ssh -N -L 8001:127.0.0.1:8001 -p <port> <user>@<gpu-server>The first request loads the checkpoint; later requests reuse the resident model. GPU requests are serialized to prevent concurrent jobs from competing for VRAM. See demo/README.md for endpoint details and upload limits.
Install the lightweight analysis package:
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev,report]"
pytest -qValidate every checked-in result and regenerate the summaries:
python scripts/validate_artifacts.py
python scripts/summarize_results.py
python scripts/analyze_research_results.py
python scripts/evaluate_optimized_router.pyFull GPU inference requires the original Stage-1 environment, MindCube images, and released model checkpoint. Example commands and frozen settings are documented in docs/EXPERIMENT_PROTOCOL.md.
| Artifact | Included | Location |
|---|---|---|
| Aggregate CSV/JSON results and statistical tests | ✅ | results/summary |
| Publication-ready charts | ✅ | assets |
| Curated demo traces and sample views | ✅ | demo/data, demo/assets |
| Architecture diagram | ✅ | assets/adaptiveplus_architecture.svg |
| Raw benchmark images and model weights | ❌ | Download from the original providers |
| Per-sample raw model generations | ❌ | Excluded from the public Git history |
Demo sample images originate from MindCube and remain subject to the original dataset terms. Raw MindCube data and model weights are intentionally not redistributed to keep the repository small and licensing boundaries clear.
src/thinker_r/ adaptive policy, routing, interventions, metrics
demo/ replay UI and FastAPI live-GPU backend
scripts/ inference, analysis, timing, and artifact validation
tests_r/ focused tests for the added research code
results/summary/ compact machine-readable evidence
reports/ human-readable results and audit report
assets/ architecture and result figures
3dthinker/ attributed upstream 3DThinker implementation
- Evaluated on one released 3B Stage-1 checkpoint and MindCube-Tiny.
- Adaptive+ is exploratory; its apparent accuracy improvement needs fresh confirmatory data.
- Answer-token confidence is relative over A–D forms, not a calibrated correctness probability.
- Fixed-duration latent interventions did not show a uniform accuracy collapse; multi-seed replication is still needed for random controls.
- This repository contributes inference, evaluation, and demo engineering—not a newly trained foundation model.
AdaSpatial-MLLM is an independent research extension of zhangquanchen/3DThinker. The original model, training code, and scientific contribution belong to Zhangquan Chen, Manyuan Zhang, Xinlei Yu, Xufang Luo, Mingze Sun, Zihao Pan, Yan Feng, Peng Pei, Xunliang Cai, and Ruqi Huang.
If you use the base model or code, cite the original 3DThinker work and follow its model/data licenses. This derivative repository retains the upstream Apache-2.0 license.
Licensed under the Apache License 2.0.






