Skip to content

TabFM (②): self-repair preprocess + sweep/plot wiring - #32

Merged
xiaohan2012 merged 4 commits into
mainfrom
tabfm-selfrepair
Jul 26, 2026
Merged

xiaohan2012 merged 4 commits into
mainfrom
tabfm-selfrepair

Conversation

@xiaohan2012

Copy link
Copy Markdown
Owner

Adds TabFM to the exp6 self-repair sweep as the 4th backbone. Follows #30 (now merged); reopened against main after #30 squash-merged (original #31 auto-closed when its base branch was deleted).

What

  • tabfm_preprocess — one clean (no-ensemble) member of TabFM's shared preprocessing: ordinal-encode categoricals + mean-impute numerics, then TabFM's own PreprocessingPipeline with the power member (CustomStandardScaler → yeo-johnson → log-based OutlierRemover). Ensemble views + the cat_mask embedding path are deferred, matching the single-clean-forward reproduction used for LimiX / TabICL / Mitra.
  • Vendored (Apache-2.0, byte-for-byte) the numeric-normalization subset of classifier_and_regressor.py → vendor/tabfm/preprocess.py (the pipeline lives in a 3942-line un-vendored estimator file; same rationale as the LimiX preprocess vendoring).
  • Wire tabfm into run_self_repair_sweep MODELS + plot_self_repair labels.
  • configs/tabfm.yaml tuned for the 24GB box: micro_batch 8, n_jobs 1 (loky+CUDA crashes), max_seq_len 512 (row-attn is O(seq²)), prior_batch_size 128 (per-step readout buffer would OOM the 64GB cgroup at 512), max_steps 100.

Verified end-to-end (GPU)

  • Finetune: 100 steps, loss 1.68 → 0.83, 25 per-depth decoders.
  • Sweep: 15/15 TabArena binary tasks, final-layer AUC ≈ native everywhere.
  • Figure 8: strong early-layer self-repair (ablate layer 1 → 0.50 chance, fully repaired by layers 2–4). Four-model comparison done.

Test

tabfm_preprocess shapes/dtype/finite on a mixed-type table (incl. a big-scale column). Passes.

🤖 Generated with Claude Code

https://claude.ai/code/session_011HRBX8LFAsrTzsyLs9M4ad

xiaohan2012 and others added 4 commits July 26, 2026 21:14
Adds TabFM to the exp6 self-repair sweep (the 4th backbone).

- tabfm_preprocess: reproduces one clean (no-ensemble) member of TabFM's
  shared preprocessing — ordinal-encode categoricals + mean-impute numerics,
  then TabFM's own PreprocessingPipeline with the `power` normalization member
  (CustomStandardScaler -> yeo-johnson -> log-based OutlierRemover). Ensemble
  views and the categorical embedding path (cat_mask) are deferred, matching
  the single-clean-forward reproduction used for the other models.
- Vendored (Apache-2.0, byte-for-byte) the numeric-normalization subset of
  classifier_and_regressor.py into vendor/tabfm/preprocess.py (the pipeline is
  in a 3942-line un-vendored estimator file, so a subset extract keeps the
  in-distribution scaling exact — like the LimiX preprocess vendoring).
- Wire `tabfm` into run_self_repair_sweep MODELS + plot_self_repair labels.

Test: tabfm_preprocess shapes/dtype/finite on a mixed-type table (incl. a
big-scale column the scaler must tame).

PR2 of 2 (GPU finetune of the 25 per-depth decoders + full sweep + Figure 8 +
four-model comparison run separately; weights are gitignored).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011HRBX8LFAsrTzsyLs9M4ad
Set the 24GB-card profile for the actual run: micro_batch_size 8 (peak ~10GB at
seq 1024) and n_jobs 1 — loky prefetch workers + the model's CUDA context crash
the run on this box, and the forward is the bottleneck so serial prior gen costs
little. save_every 50 so a long run survives a stall.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011HRBX8LFAsrTzsyLs9M4ad
TabFM's row attention is O(seq^2), so seq 1024 runs minutes/step (~16h for
200). Drop to 512 (like Mitra) -> ~70s/step, ~4h. Decoder decodability doesn't
need the full 1024-row context.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011HRBX8LFAsrTzsyLs9M4ad
- prior_batch_size 512 -> 128: each step parks all 25 depths' test-row readouts
  for prior_batch_size tables in host RAM; TabFM's large icl_dim blew the ~64GB
  cgroup on step 2 (SIGKILL 137). 128 keeps the working set ~1/4.
- max_steps 200 -> 100: overkill for these linear decoders; halves runtime.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011HRBX8LFAsrTzsyLs9M4ad
@xiaohan2012
xiaohan2012 merged commit 841afe1 into main Jul 26, 2026
1 check passed
@xiaohan2012
xiaohan2012 deleted the tabfm-selfrepair branch July 26, 2026 18:15
@xiaohan2012

Copy link
Copy Markdown
Owner Author

The four-model comparison this PR completes

TabFM was the fourth and last backbone, so this is where the cross-architecture picture closes. The figure lived under a gitignored out/; committing it here for the record (analysis/exp6-figures/, pinned).

Setup: 15 TabArena binary tasks per model · natural OpenML split · identity-skip each layer, decode every depth, normalize per task by native final AUC (floored 0.5), then average. One measurement path, four adapters.

Self-repair across four TFMs

Black = no ablation. Colored = trajectory from each skip point, dark→light by ablated layer. Red ✗ = immediate post-skip reading.

model layers unrecoverable late layers
LimiX-2M 12 L0–L2, still 0.05–0.08 short at the output back on baseline within a layer or two
Mitra 12 L0/L1 fold back
TabICLv2 12 essentially none (baseline already starts at 0.95) one mid-stack dip that recovers
TabFM 24 L1 only (drops to 0.50 chance, repaired by L2–L4) fold back

The drop-then-recover shape and the early-layer exception hold in all four. The adapter contract is what makes this comparison meaningful: identical capture, intervention, and decode path across a 3D single-stream model (TabICL, TabFM), a 4D model (LimiX), and a double-stream one (Mitra) — the only per-model code is the adapter and the preprocessing.

Caveat worth recording

The y-axis is AUC, which is rank-based and ceiling-blind: once the classes are separated, later layers can keep doing signed work without moving the curve. That is why the late-layer skips look so flat. A magnitude coordinate makes dips visible that AUC hides entirely — see #39.

And all of it is a total effect: the ablation propagates freely downstream at every point of measurement, so a curve that dips and returns is equally consistent with active repair and with plain redundancy. Separating the two needs a direct effect — #45.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant