TabFM (②): self-repair preprocess + sweep/plot wiring - #32
Conversation
Adds TabFM to the exp6 self-repair sweep (the 4th backbone). - tabfm_preprocess: reproduces one clean (no-ensemble) member of TabFM's shared preprocessing — ordinal-encode categoricals + mean-impute numerics, then TabFM's own PreprocessingPipeline with the `power` normalization member (CustomStandardScaler -> yeo-johnson -> log-based OutlierRemover). Ensemble views and the categorical embedding path (cat_mask) are deferred, matching the single-clean-forward reproduction used for the other models. - Vendored (Apache-2.0, byte-for-byte) the numeric-normalization subset of classifier_and_regressor.py into vendor/tabfm/preprocess.py (the pipeline is in a 3942-line un-vendored estimator file, so a subset extract keeps the in-distribution scaling exact — like the LimiX preprocess vendoring). - Wire `tabfm` into run_self_repair_sweep MODELS + plot_self_repair labels. Test: tabfm_preprocess shapes/dtype/finite on a mixed-type table (incl. a big-scale column the scaler must tame). PR2 of 2 (GPU finetune of the 25 per-depth decoders + full sweep + Figure 8 + four-model comparison run separately; weights are gitignored). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011HRBX8LFAsrTzsyLs9M4ad
Set the 24GB-card profile for the actual run: micro_batch_size 8 (peak ~10GB at seq 1024) and n_jobs 1 — loky prefetch workers + the model's CUDA context crash the run on this box, and the forward is the bottleneck so serial prior gen costs little. save_every 50 so a long run survives a stall. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011HRBX8LFAsrTzsyLs9M4ad
TabFM's row attention is O(seq^2), so seq 1024 runs minutes/step (~16h for 200). Drop to 512 (like Mitra) -> ~70s/step, ~4h. Decoder decodability doesn't need the full 1024-row context. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011HRBX8LFAsrTzsyLs9M4ad
- prior_batch_size 512 -> 128: each step parks all 25 depths' test-row readouts for prior_batch_size tables in host RAM; TabFM's large icl_dim blew the ~64GB cgroup on step 2 (SIGKILL 137). 128 keeps the working set ~1/4. - max_steps 200 -> 100: overkill for these linear decoders; halves runtime. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011HRBX8LFAsrTzsyLs9M4ad
The four-model comparison this PR completesTabFM was the fourth and last backbone, so this is where the cross-architecture picture closes. The figure lived under a gitignored Setup: 15 TabArena binary tasks per model · natural OpenML split · identity-skip each layer, decode every depth, normalize per task by native final AUC (floored 0.5), then average. One measurement path, four adapters. Black = no ablation. Colored = trajectory from each skip point, dark→light by ablated layer. Red ✗ = immediate post-skip reading.
The drop-then-recover shape and the early-layer exception hold in all four. The adapter contract is what makes this comparison meaningful: identical capture, intervention, and decode path across a 3D single-stream model (TabICL, TabFM), a 4D model (LimiX), and a double-stream one (Mitra) — the only per-model code is the adapter and the preprocessing. Caveat worth recordingThe y-axis is AUC, which is rank-based and ceiling-blind: once the classes are separated, later layers can keep doing signed work without moving the curve. That is why the late-layer skips look so flat. A magnitude coordinate makes dips visible that AUC hides entirely — see #39. And all of it is a total effect: the ablation propagates freely downstream at every point of measurement, so a curve that dips and returns is equally consistent with active repair and with plain redundancy. Separating the two needs a direct effect — #45. |

Adds TabFM to the exp6 self-repair sweep as the 4th backbone. Follows #30 (now merged); reopened against
mainafter #30 squash-merged (original #31 auto-closed when its base branch was deleted).What
tabfm_preprocess— one clean (no-ensemble) member of TabFM's shared preprocessing: ordinal-encode categoricals + mean-impute numerics, then TabFM's ownPreprocessingPipelinewith thepowermember (CustomStandardScaler → yeo-johnson → log-based OutlierRemover). Ensemble views + thecat_maskembedding path are deferred, matching the single-clean-forward reproduction used for LimiX / TabICL / Mitra.classifier_and_regressor.py→vendor/tabfm/preprocess.py(the pipeline lives in a 3942-line un-vendored estimator file; same rationale as the LimiX preprocess vendoring).tabfmintorun_self_repair_sweepMODELS+plot_self_repairlabels.configs/tabfm.yamltuned for the 24GB box:micro_batch 8,n_jobs 1(loky+CUDA crashes),max_seq_len 512(row-attn is O(seq²)),prior_batch_size 128(per-step readout buffer would OOM the 64GB cgroup at 512),max_steps 100.Verified end-to-end (GPU)
Test
tabfm_preprocessshapes/dtype/finite on a mixed-type table (incl. a big-scale column). Passes.🤖 Generated with Claude Code
https://claude.ai/code/session_011HRBX8LFAsrTzsyLs9M4ad