Skip to content

Latest commit

 

History

History
250 lines (210 loc) · 10.7 KB

File metadata and controls

250 lines (210 loc) · 10.7 KB

Remote implementation status

Updated: 2026-08-17

Host and storage

  • Host: autodl-container-9cjes0s087-77ab373c
  • GPU: NVIDIA GeForce RTX 5090, compute capability 12.0
  • Project root: /root/autodl-tmp/3Dyujia
  • Data disk: 150 GB total, 129 GB free after the independent-audit stage
  • System disk: 30 GB total, 2.0 GB used

Verified environment

  • Environment: /root/autodl-tmp/3Dyujia/envs/hmr2
  • PyTorch: 2.7.1+cu128
  • torchvision: 0.22.1+cu128
  • NumPy: 1.26.4
  • PyTorch Lightning: 1.8.6
  • CUDA matrix smoke test: passed on RTX 5090
  • EGL offscreen renderer: passed
  • Project tests: 48 passed on the server and 48 passed locally

HMR baseline

  • Upstream: official 4DHumans repository
  • Commit: efe18deff163b29dff87ddbd575fa29b716a356c
  • Source: /root/autodl-tmp/3Dyujia/third_party/4D-Humans
  • Official model archive SHA-256: 0fdf9e66ec97503fe1b995f4942e021a6df748f4ccbc5574718742d975a6e19b
  • Checkpoint: /root/autodl-tmp/3Dyujia/models/4dhumans/logs/train/multiruns/hmr2/0/checkpoints/epoch=35-step=1000000.ckpt
  • Checkpoint format audit: valid ZIP archive with 578 entries

The official archive is an uncompressed POSIX tar despite its .tar.gz filename. It was extracted according to the detected format after matching local and remote hashes.

Implemented project components

  • explicit bbox single-image HMR runner without Detectron2;
  • HMR feature NPZ and metadata output;
  • gravity-preserving 3D-joint canonicalization;
  • SMPL rotation-matrix to 6D features;
  • hierarchical 6/20/82 MLP classifier;
  • cached-feature dataset and training entrypoint;
  • official-split-to-CSV manifest adapter;
  • resumable, validated Yoga-82 URL downloader;
  • deterministic, non-destructive Yoga-82 cleaning and exact-leakage audit;
  • deterministic 82-class HMR audit sampler;
  • resumable single-load HMR audit runner;
  • COCO Faster R-CNN person-detection gate;
  • data-disk cache and storage guard scripts.

Licensed SMPL model

  • Preserved licensed Python 2 source: assets/smpl/SMPL_NEUTRAL_v1.1.0_py2_original.pkl
  • Converted Python 3 runtime model: models/4dhumans/data/smpl/SMPL_NEUTRAL.pkl
  • Runtime SMPL smoke test: 6,890 finite vertices, 45 output joints, and 13,776 faces.

Yoga-82 download

  • Official metadata: 28,450 URLs, 20,994 train and 7,456 test rows;
  • hierarchy verified as 6 / 20 / 82 classes;
  • archive SHA-256: 345d646a65c0dfa6ca0c205a21762f5ae12072bcd636fabda02733292e0985a9;
  • full first-pass download completed all 28,450 rows;
  • 11,821 files are locally available and 16,629 legacy URLs failed;
  • final summary: data/raw/yoga82/logs/full_pass1.summary.json.

Yoga-82 clean v1

  • Frozen clean manifest: 11,760 unique images;
  • train/test: 8,727 / 3,033;
  • complete 6 / 20 / 82 coverage in both splits;
  • 0 exact SHA-256 overlaps between train and test;
  • 61 reachable files excluded: 30 redundant exact copies, 21 files in label-conflicting duplicate groups, 9 uniform/blank images, and 1 image below the 64-pixel minimum dimension;
  • all 11,821 original downloaded files remain untouched;
  • manifest and audit reports: data/processed/yoga82/manifests/.

Full frozen-HMR extraction v1

  • Input: all 11,760 rows from yoga82_reachable_clean_v1.csv;
  • Faster R-CNN detector: 10,623 detected, 1,137 explicit no-person rejections, and 0 decode errors after adding a Pillow first-frame fallback for 47 GIF files stored with .jpg names;
  • frozen HMR2: 10,623 / 10,623 detected rows processed, 0 HMR failures;
  • successful train/test features: 7,622 / 3,001;
  • every successful cache has finite SMPL pose, 3D joints, camera parameters, and an explicit 216-dimensional classifier vector;
  • successful train and test rows both retain complete 6 / 20 / 82 coverage;
  • mesh vertices and full-dataset renders were intentionally omitted;
  • independent validation passed for 10,623 NPZ and 10,623 metadata files;
  • output size: 96 MB;
  • detector output: outputs/detection/yoga82_clean_v1/;
  • HMR output: outputs/hmr2/yoga82_clean_v1/.

HMR audit v1

  • Frozen audit: 164 images, two per fine class, complete 6 / 20 / 82 coverage;
  • full-frame execution: 164 / 164 numerical outputs, all finite;
  • person detector: 147 accepted and 17 explicit no-person rejections;
  • detected-box HMR: 147 finite outputs with 6 / 20 / 81 coverage;
  • full evidence boundary and seed visual review: docs/HMR_AUDIT_V1.md.

The numerical success count is not a visual 3D-accuracy result. Yoga-82 has no ground-truth SMPL, and strong self-contact, illustrations, small people, and suspected label noise remain visible failure modes.

SMPL/3D MLP baseline v1

  • accepted data: 7,622 train and 3,001 test feature caches;
  • fine-class-stratified fit/validation split: 6,480 / 1,142 per run;
  • three formal seeds: 2026, 2027, and 2028;
  • 82-class test Top-1: 91.44 ± 0.47%;
  • 82-class test Top-5: 97.96 ± 0.05%;
  • 82-class test Macro-F1: 90.08 ± 0.51%;
  • accepted-test coverage: 3,001 / 3,033 = 98.945%;
  • 82-class Top-1 with rejected test rows counted as errors: 90.47%;
  • all checkpoints, test predictions, per-class recall, and confusion matrices passed the aggregate consistency validation;
  • report: docs/SMPL3D_MLP_V1.md;
  • artifacts: experiments/smpl3d_mlp_v1/.

This result establishes a functioning 3D baseline only. It is not evidence of a 3D advantage until matched 2D-keypoint and RGB baselines are complete.

Matched 2D-keypoint MLP baseline v1

  • COCO-17 Keypoint R-CNN inference uses the exact frozen HMR person boxes;
  • 10,623 / 10,623 HMR-accepted rows produced valid 51D 2D features;
  • feature extraction took 251.9 seconds and introduced no new rejections;
  • three formal seeds: 2026, 2027, and 2028;
  • 82-class test Top-1: 71.02 ± 0.48%;
  • 82-class test Top-5: 90.65 ± 0.12%;
  • 82-class test Macro-F1: 67.15 ± 0.69%;
  • SMPL/3D exceeds this 2D baseline by 20.42 Top-1 percentage points and 22.93 Macro-F1 percentage points at 82 classes;
  • prediction identities and targets match row-for-row across both baselines;
  • report: docs/POSE2D_MLP_V1.md;
  • features: outputs/pose2d/yoga82_hmr_matched_v1/;
  • artifacts: experiments/pose2d_mlp_v1/.

The difference does not isolate depth because HMR2 and Keypoint R-CNN use different estimators.

Same-HMR projected-2D ablation v1

  • exact same 24 HMR joints, accepted rows, splits, and three seeds as SMPL/3D;
  • HMR2 perspective projection produces a 48D (x, y) representation;
  • 10,623 / 10,623 caches valid; generation took 11.8 seconds;
  • 82-class Top-1: 88.89 ± 0.22%;
  • 82-class Macro-F1: 86.93 ± 0.18%;
  • full SMPL/3D improvement: +2.54 Top-1 and +3.14 Macro-F1 percentage points;
  • projected HMR improvement over independent COCO-17: +17.87 Top-1 points;
  • report: docs/HMR_PROJECTED2D_MLP_V1.md;
  • features: outputs/hmr_projected2d/yoga82_hmr_matched_v1/;
  • artifacts: experiments/hmr_projected2d_mlp_v1/.

The remaining full-versus-projected gap combines depth, 3D normalization, and SMPL rotations; it is not a pure depth measurement.

HMR feature decomposition v1

  • exact split of the original 216D vector into 72D canonicalized 3D joints and 144D SMPL rotations;
  • both caches contain 10,623 / 10,623 valid matched samples;
  • concatenating every new 72D and 144D vector exactly reproduces its original 216D vector;
  • 82-class Top-1: projected 2D 88.89%, joints 88.91%, rotations 90.39%, full fusion 91.44%;
  • 82-class Macro-F1: projected 2D 86.93%, joints 87.30%, rotations 89.08%, full fusion 90.08%;
  • rotations are the strongest individual block; full fusion adds another 1.04 Top-1 points over rotations-only;
  • report: docs/HMR_FEATURE_ABLATION_V1.md;
  • suite artifact: experiments/hmr_feature_ablation_v1.json.

Clean-v2 experiments and HMR quality triage

  • clean-v2 manifest: 11,605 images, including 8,575 train and 3,030 test;
  • accepted HMR features: 10,474, including 7,473 train and 3,001 test;
  • exact cross-split leakage removed without deleting raw files;
  • 82-class full-3D Top-1: 91.05%, Macro-F1: 89.81%;
  • fine-tuned RGB Top-1: 82.59%, Macro-F1: 79.70%;
  • validation-selected RGB + full-3D late fusion Top-1: 92.87%, Macro-F1: 92.06%;
  • independent HMR visual-quality model: 145 audited samples, OOF ROC-AUC 0.698 and average precision 0.779;
  • automatic score-based rejection is disabled because its audit precision did not reach the predeclared 75% target;
  • HMR quality accept/review test strata obtain 92.64% / 87.98% downstream full-3D classification accuracy;
  • complete class recall, confusion analysis, vector paper figures, and the success/failure gallery are in experiments/clean_v2/;
  • report: docs/HMR_QUALITY_AND_RESULTS_STAGE_REPORT.md.

Current work

  1. collect external labels for the completed 328-image blinded HMR pack and compute inter-rater agreement;
  2. adjudicate the frozen 240-pair perceptual near-duplicate extension;
  3. decide whether those decisions require clean-v3 and a retraining cycle;
  4. retrieve and verify literature for Introduction and Related Work;
  5. optionally add MoYo/SMPL-X ground-truth 3D validation;
  6. keep webcam integration deferred unless a live demonstration is later required.

Independent 2D benchmark extension (clean-v2)

  • independent estimators: Keypoint R-CNN ResNet-50-FPN, MediaPipe BlazePose heavy, and ViTPose-Base simple;
  • Keypoint R-CNN and ViTPose produced 10,474 / 10,474 matched HMR-accepted features;
  • MediaPipe produced 10,263 / 10,474 features, with 211 no-pose failures;
  • strict five-method common set: 7,311 train and 2,952 test images, or 97.43% of the complete clean-v2 test split;
  • three-seed 82-class Top-1 on the common set: Keypoint R-CNN 71.85%, MediaPipe 84.65%, ViTPose 80.76%, same-HMR projected 2D 87.82%, and full SMPL/3D 91.52%;
  • full SMPL/3D gains over those four 2D baselines: +19.67, +6.87, +10.76, and +3.70 Top-1 percentage points;
  • all methods use identical row identities, targets, MLP protocol, and seeds;
  • report: docs/POSE2D_BENCHMARKS_CLEAN_V2.md;
  • artifacts: experiments/pose2d_benchmarks_matched_clean_v2/;
  • server feature caches: outputs/pose2d_benchmarks/;
  • automated tests: 53 / 53 passed.

Independent validation and paper consolidation

  • blinded HMR validation selection: 328 images, four per fine class;
  • selection balance: 164 quality-model accept, 164 review;
  • previous 164-image seed audit excluded in full;
  • HMR2 render completion: 328 / 328;
  • review assets: 656 files plus frozen annotation protocol and browser index;
  • independent external labels: pending and not reported as completed;
  • perceptual duplicate extension: 240 prioritized pairs in 40 review pages;
  • duplicate ranker OOF ROC-AUC / AP: 0.836 / 0.851, for prioritization only;
  • consolidated paper tables: six methods in Markdown, CSV, and LaTeX;
  • paper artifacts: evidence map, blueprint, Methods/Results draft, and pre-submission audit;
  • stage report: docs/INDEPENDENT_VALIDATION_AND_PAPER_STAGE.md.