Block 8 (exp6, ⑤): self-repair repro driver + Figure-8 plot - #18
Conversation
- load_decoders(path, adapter) in evaluation.layerwise: shared loader for the 13 fine-tuned decoders; dedupes the inline loading in test_layerwise - scripts/run_self_repair.py: loop the 15 TabArena binary tasks -> preprocess -> ablation_sweep + native_final_auc [+ ablation_diffs]; CLI for subsample sizes (0 = all rows), output path, and --skip-diffs; per-task try/except so one bad table can't sink the run - scripts/plot_self_repair.py: Figure-8 self-repair trajectories, per-dataset normalized by native final AUC then averaged; black baseline + colored per-ablation trajectories + red-x immediate-drop connectors - viz dependency-group (matplotlib); gitignore out/ Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X7n7diu4dvRuyRKgWeGnWK
…ad_decoders test - rename scripts/run_self_repair.py -> run_self_repair_sweep.py (pairs with plot_self_repair.py; "sweep" names what it does), fix docstring refs - test_load_decoders_returns_one_per_depth: round-trip a toy template through a tmp dir instead of requiring trained weights, so it runs unskipped in CI — it's pure plumbing (read N state dicts -> N decoders), not a numeric check Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X7n7diu4dvRuyRKgWeGnWK
Per #18 review: the old test asserted the fine-tuned decoders "build up with depth" with a hard-coded aucs[-1] > aucs[0] + 0.2. That's a post-hoc threshold 追认ing an observed scientific result, and a behavioral claim rather than a functional check — it doesn't belong in the unit suite (and the logit-lens story is already shown by the Figure-8 reproduction). Recast it as the one thing no other test covers: a smoke test that the loaded fine-tuned weights plug into the eval pipeline on real LimiX and decode signal at the final layer (finite AUCs in [0,1], final > chance). Dropped the brittle depth-climb and >0.8 asserts. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X7n7diu4dvRuyRKgWeGnWK
The real-LimiX test asserted the fine-tuned decoders "build up with depth" (aucs[-1] > aucs[0] + 0.2): a post-hoc threshold on an observed scientific result, not a functional check. Once that numeric assert goes, the test needs no trained weights and only re-checks shape/range — already covered by the toy test_predict_layers_returns_valid_probs_per_depth and, on real LimiX, by test_evaluation_real. So remove it rather than keep a redundant shell; its sole-use imports (pytest, Path, LimixAdapter) and TRAINED_DECODERS go with it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X7n7diu4dvRuyRKgWeGnWK
The self-repair repro driver (#18) produced its figures under a gitignored out/, so the PR carries no plot. Commit the two that document the result: - pr18_limix_2m_fig8.png - the driver's own output (LimiX-2M, 15 tasks) - pr18_all4_fig8.png - the same figure for all four TFMs Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Results — the figures this PR producedThe driver's output landed under a gitignored Setup: LimiX-2M · 15 TabArena binary tasks · no subsample (natural OpenML split) · skip-ablate each layer, decode every depth, normalize per task by native final AUC (floored 0.5), then average. LimiX-2M — the driver's own outputBlack = no ablation. Colored = trajectory from each skip point, dark→light by ablated layer. Red ✗ = immediate post-skip reading, dashed line to where the baseline was.
All four modelsSame figure for the other three (added after this PR, same driver and settings):
Caveat worth recordingThe recovery here is measured on AUC, which is rank-based and ceiling-blind — once the classes are separated, later layers can keep doing signed work without moving the curve. That is why the late-layer skips look so flat above. Switching the y-axis to a magnitude metric (margin / GT-logit) makes dips visible that AUC hides entirely — see #39. More importantly, all of this is a total effect: the ablation propagates freely downstream at every point of measurement. A trajectory that dips and returns to baseline is equally consistent with active repair and with plain redundancy. Separating the two needs a direct effect, which is #37/#45. |


Ties exp6 together end-to-end: a driver that runs the self-repair sweep across the 15 TabArena binary tasks, plus the Figure-8 trajectory plot. Everything downstream (sweep / metrics / eval) already landed in #16 / #17; this is the thin orchestration + visualization layer on top.
What's here
scripts/run_self_repair.py— per-dataset: load → (optional) subsample → LimiX preprocess →ablation_sweep+native_final_auc[+ablation_diffs]. Dumps JSON.--subsample-train/-test(0 = full),--skip-diffs,--out. Per-dataset try/except so one bad table can't sink the run.scripts/plot_self_repair.py— paper Figure 8: black baseline + per-ablation trajectories from the skip point + red-x immediate-drop connectors, normalized per-dataset by native final AUC (floored 0.5) then averaged.evaluation/layerwise.py::load_decoders— one fine-tuned decoder per capture depth from a weights dir (dedupes inline loading; used by the driver).pyproject.toml—vizdependency-group (matplotlib), kept out of core/CI..gitignore—weights/,out/, run logs (trained decoders stay out of version control).Verification
Full run (15/15 tasks, no subsample) reproduces the self-repair shape: black baseline climbs 0.59→1.0; early-layer ablations show a large immediate drop then recovery; late-layer ablations barely dent it.
Figure artifacts live under gitignored
out/(not committed).