This demo has two modes:
- Replay is dependency-free and replays real per-sample traces from the checked-in Stage-1 evaluation artifacts. It does not fabricate outputs or require a GPU.
- Live GPU accepts 2–4 user-supplied views and compares fixed-12 with the Adaptive+ policy. Adaptive+ routes by task, records A–D answer-token confidence, and adds fixed-4/fixed-8 voting for difficult Among samples. An opt-in left/right mirror consistency check is also available.
Representative real-GPU captures are stored in assets/screenshots. The
two-view cascade sample and its images are stored in assets/sample-cascade.
From the repository root:
python -m http.server 7860 --directory demoOpen http://localhost:7860.
Use the same Python environment that already runs the patched 3DThinker Stage-1 model, then add the small web extra:
pip install -e ".[demo]"
THINKER_MODEL_PATH=/path/to/3DThinker-Mindcube \
python -m uvicorn demo.live_server:app --host 127.0.0.1 --port 8001The API intentionally binds to loopback. When the GPU is on a remote server, forward it to the machine serving the browser:
ssh -N -L 8001:127.0.0.1:8001 -p <port> <user>@<server>Then switch the page from REPLAY to LIVE GPU. The first request loads the checkpoint; later
requests reuse the resident model. GPU jobs are serialized so concurrent browser requests cannot
compete for the same device.
API endpoints:
GET /healthreports GPU and model residency.POST /infer/compareruns fixed-12 plus Adaptive+ for the research comparison view.POST /infer/optimizedruns Adaptive+ only, avoiding the fixed-12 display baseline in deployment.
Both POST endpoints accept multipart images, a question, four choices, an optional reference
answer, task_type (auto, among, rotation, or around), and optional symmetry_check=true.
Uploads are limited to 2–4 JPEG/PNG/WebP files, 10 MB each, and resized to at most 640 px before
inference. Browser access is restricted by THINKER_ALLOWED_ORIGINS, which defaults to the local
demo origins on port 7860.
python scripts/build_demo_data.pyThe builder joins the frozen MindCube question records with the fixed-4, fixed-8, fixed-12, and
adaptive JSONL outputs, then writes demo/data/demo-data.json. The four displayed images are the
actual inputs for the curated coffee-table scene.
- Multi-view 3D spatial reasoning input.
- A 12-step fixed latent budget versus per-sample adaptive early exit.
- The hidden-state cosine-distance trajectory and the frozen
epsilon=0.04,patience=2rule. - Per-sample predictions, correctness, recorded latency, and exit reason.
- Full 950-sample holdout results, so a curated example is not presented as aggregate evidence.
- The exploratory Adaptive+ result (68.53%, 651/950) and its non-significant paired tests are kept separate from the confirmatory adaptive early-exit claim.
The Replay and Live modes use the same result cards and latent oscilloscope, making recorded benchmark evidence and fresh user inputs directly comparable without mixing their provenance.
