Skip to content

Latest commit

 

History

History
112 lines (87 loc) · 6.23 KB

File metadata and controls

112 lines (87 loc) · 6.23 KB

AdaSpatial-MLLM: training-free adaptive latent inference

Scope

This project evaluates and improves the released 3DThinker Stage-1 checkpoint without training or claiming a new foundation model. The research contribution is a reproducible inference and evaluation system plus controlled evidence about latent computation.

The implementation starts from upstream commit c9469e01b719310b0eaecc1133317e4ecfc74d8c. All changes live on codex/repro-adaptive.

Research questions

  1. Does the released checkpoint run end to end on current RTX 5090 software?
  2. How does a fixed latent budget of 4, 8, or 12 affect accuracy and cost?
  3. Can hidden-state convergence reduce the mean latent budget without a material accuracy loss?
  4. Do zero, distribution-matched random, and feature-shuffled latent states change predictions?
  5. Can a post-hoc, training-free task router and hard-example cascade improve answer accuracy?

Assets and integrity

  • Model: jankin123/3DThinker-Mindcube, Stage-1 Qwen2.5-VL checkpoint.
  • Model repository revision: 69a70411605f86ec69bada0a625bb96ddee995d9.
  • Checkpoint-declared latent size: 12.
  • Tensor payload declared by the safetensors index: 7,649,171,816 bytes.
  • Dataset: MLL-Lab/MindCube, MindCube_tinybench.jsonl, 1,050 questions.
  • Dataset archive SHA-256: 714c8961d66e9662c826d2202493a7d86cbe31797cf41e3a0567cb4d72b0aaa8.
  • Tiny benchmark composition: 600 among, 200 rotation, and 250 around questions.

The downloaded MindCube archive passes unzip -t. Model loading is forced offline during experiments so a run cannot silently switch checkpoints.

Evaluation design

Primary decoding is greedy and deterministic with seed 42. A secondary reproduction run uses the author's sampling settings (temperature=0.7, top_p=0.9) and must be reported separately. The formal maximum answer budget is 1,024 tokens; any sequence reaching it is marked truncated. Images whose longest side exceeds 640 pixels are proportionally resized with Lanczos. This activates the resize helper present in the author's inference script and is required for the public 1296x968 rotation images to fit the 32 GB GPU; smaller among and around images are unchanged.

The conditions are:

Family Conditions Purpose
Fixed compute 4, 8, and 12 latent pads Accuracy/cost curve
Adaptive minimum 4, maximum 12, cosine convergence Dynamic compute
Causal controls normal, zero, matched-random, shuffled features Test whether latent content matters

The later Adaptive+ prototype is deliberately separated from the confirmatory matrix. It routes Among questions to adaptive inference, Rotation/Around to fixed-12, and escalates non-converged Among samples to a fixed-4/fixed-8/adaptive majority vote. Since this rule was designed after examining holdout task breakdowns, all Adaptive+ numbers are labeled exploratory. Answer-token confidence and opt-in horizontal-mirror consistency are live diagnostics and are not included in the frozen-matrix replay result.

shuffle_features applies one deterministic permutation to the hidden dimension at every latent step. This preserves the exact per-step multiset of values while disrupting learned feature semantics. It is deliberately named more precisely than a generic "shuffle" intervention.

The initial exploratory controls allowed the checkpoint to emit its latent-end token. Because corrupted states changed that stopping behavior, they confounded latent content with latent compute. The confirmatory controls therefore use --force-latent-steps: all four conditions execute exactly 12 latent pads on the same 200 held-out IDs. Each corrupted run also records cosine similarity and relative L2 distance between the original and injected states, providing a manipulation check.

The convergence threshold is selected from a calibration slice using fixed-12 hidden-state traces. The final comparison is performed on held-out IDs. Threshold selection and final evaluation IDs must be recorded in the result manifest.

Measurements

Every result row records the prediction, ground truth, correctness, full response, input and output token counts, truncation, CUDA wall time, peak allocated VRAM, latent pads, exit reason, and cosine distance trajectory. Report:

  • overall and task-level accuracy;
  • Wilson 95% confidence intervals;
  • paired improvements/regressions and exact McNemar p-value;
  • mean latent pads and early-exit fraction;
  • clean batch-1 latency and peak VRAM;
  • batch-2 throughput as a separate systems measurement.

Accuracy sweeps may run as two concurrent batch-1 processes only after the selected category's vision-memory peak has been measured. This is safe for the initial among slice but can OOM on larger rotation inputs. Their latency fields are never used. Timing runs execute alone after a warm-up and are labeled separately.

The clean timing set contains 50 held-out examples: 20 among, 15 rotation, and 15 around. Each condition runs sequentially in a fresh process. Three additional examples, one per task, run first as GPU warm-up and are excluded from all timing statistics. Report mean, median, p95, standard deviation, output-token count, mean latent steps, and peak allocated VRAM.

Confirmed upstream issues

  1. The original generation loop hard-codes eight latent pads although the released checkpoint declares twelve.
  2. Latent token IDs are hard-coded instead of read from checkpoint configuration.
  3. The released tokenizer owns the chat template, but the original script calls the processor's missing template and fails before inference.
  4. The original script hard-codes paths and does not save structured predictions or runtime metrics.
  5. Upstream batch latent state is globally synchronized and is unsuitable for true per-sample adaptive exit. Adaptive evaluation therefore uses batch 1.

Completion criteria

  • The 100-example calibration split and 950-example held-out split remain disjoint.
  • All 950 held-out examples complete for fixed 4/8/12 and the selected adaptive policy.
  • All three causal controls complete on the same paired IDs.
  • No unreported truncation or failed sample remains.
  • Raw JSONL, manifests, aggregate tables, plots, and exact commands are retained.
  • Claims in the CV and repository README match measured results and clearly state "training-free".