Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 
 
 
 
 
 
 

README.md

AdaSpatial-MLLM Interactive Demo

This demo has two modes:

  • Replay is dependency-free and replays real per-sample traces from the checked-in Stage-1 evaluation artifacts. It does not fabricate outputs or require a GPU.
  • Live GPU accepts 2–4 user-supplied views and compares fixed-12 with the Adaptive+ policy. Adaptive+ routes by task, records A–D answer-token confidence, and adds fixed-4/fixed-8 voting for difficult Among samples. An opt-in left/right mirror consistency check is also available.

AdaSpatial-MLLM demo preview

Representative real-GPU captures are stored in assets/screenshots. The two-view cascade sample and its images are stored in assets/sample-cascade.

Run locally

From the repository root:

python -m http.server 7860 --directory demo

Open http://localhost:7860.

Enable Live GPU mode

Use the same Python environment that already runs the patched 3DThinker Stage-1 model, then add the small web extra:

pip install -e ".[demo]"
THINKER_MODEL_PATH=/path/to/3DThinker-Mindcube \
  python -m uvicorn demo.live_server:app --host 127.0.0.1 --port 8001

The API intentionally binds to loopback. When the GPU is on a remote server, forward it to the machine serving the browser:

ssh -N -L 8001:127.0.0.1:8001 -p <port> <user>@<server>

Then switch the page from REPLAY to LIVE GPU. The first request loads the checkpoint; later requests reuse the resident model. GPU jobs are serialized so concurrent browser requests cannot compete for the same device.

API endpoints:

  • GET /health reports GPU and model residency.
  • POST /infer/compare runs fixed-12 plus Adaptive+ for the research comparison view.
  • POST /infer/optimized runs Adaptive+ only, avoiding the fixed-12 display baseline in deployment.

Both POST endpoints accept multipart images, a question, four choices, an optional reference answer, task_type (auto, among, rotation, or around), and optional symmetry_check=true.

Uploads are limited to 2–4 JPEG/PNG/WebP files, 10 MB each, and resized to at most 640 px before inference. Browser access is restricted by THINKER_ALLOWED_ORIGINS, which defaults to the local demo origins on port 7860.

Rebuild the demo payload

python scripts/build_demo_data.py

The builder joins the frozen MindCube question records with the fixed-4, fixed-8, fixed-12, and adaptive JSONL outputs, then writes demo/data/demo-data.json. The four displayed images are the actual inputs for the curated coffee-table scene.

What the demo communicates

  • Multi-view 3D spatial reasoning input.
  • A 12-step fixed latent budget versus per-sample adaptive early exit.
  • The hidden-state cosine-distance trajectory and the frozen epsilon=0.04, patience=2 rule.
  • Per-sample predictions, correctness, recorded latency, and exit reason.
  • Full 950-sample holdout results, so a curated example is not presented as aggregate evidence.
  • The exploratory Adaptive+ result (68.53%, 651/950) and its non-significant paired tests are kept separate from the confirmatory adaptive early-exit claim.

The Replay and Live modes use the same result cards and latent oscilloscope, making recorded benchmark evidence and fresh user inputs directly comparable without mixing their provenance.