feat: path-patching DE/TE on the native head — DE–TE scatter (#44, Part 1) - #45
Conversation
…t 1) The DE–TE scatter (Hydra Fig 2c analog) needs two rulers in one basis. Add `evaluation/direct_effect.py`: DE (direct effect, path-patching 取法 B) and TE (total effect), both read through the model's own head — not per-layer tuned decoders — so they're comparable by construction. - DE(m): residual is additive, so freezing downstream at clean is arithmetic, r_L − a_m + ã_m, off one clean forward. Faithful head (final LN σ recomputed), not a fixed û on a single layer. - TE(m): ablate-and-react (inject_delta / skip_layer), read the final residual through the same head. DE and TE share one donor draw per (m, donor). - resample ã_m (role-matched donor δ, metric averaged) reusing the #35 machinery; zero as the cross-check. Plus `run_direct_effect_sweep.py` (per-task DE/TE + native-final normalizer) and `plot_de_te_scatter.py` (D1). Additive PR: touches no existing module. Tests (tests/test_direct_effect.py): the headline invariant is that the O(1)-forward arithmetic DE reproduces a real frozen-downstream forward (3D / 4D / double-stream); plus DE == logit-lens on a linear head (取法 A == B) and the resample path's shape. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
- rename for readability: native_head_logits -> final_decoder_logits, _swap ->
patched_residual; descriptive vars (clean_contributions / clean_residual /
replacement); unify "head" -> "decoder"; document a_m = clean write r_m - r_{m-1}.
- drop zero ablation entirely (module + sweep + tests) — resample only, as agreed.
- dedup: move clone_residual to utils (was self_repair._clone_residual); factor
native_final_logits out of native_final_auc so TE reuses it instead of a bespoke
forward (_te_forward deleted).
- add type hints (Residual / Coords aliases, signatures + returns).
Tests rewritten to the primitive-level invariant with an arbitrary replacement ã_m
(no zero): patched_residual == frozen-downstream forward (3D / 4D / double-stream),
plus the linear-head 取法 A == B identity. Re-smoked on LimiX-2M end-to-end.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
|
Review pass in
Tests rewritten to the primitive invariant with an arbitrary ã_m (no zero): |
- naming: Coords -> MetricPair, _coords -> _metric_pair, _effect -> _drop_from_clean, _mean_coords -> _mean_over_donors (more specific names per review). - English: translate the remaining 取法 A/B -> method A/B across module, scripts, tests. - dedup: move the duplicated _subsample / _load_record out of both sweep scripts into evaluation/datasets.py as subsample() + load_task_record(); scripts import them. Re-smoked on LimiX-2M end-to-end; ruff + affected tests green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
Un-aggregate the ~500 test rows per (task, layer) to close the sign-cancellation loophole: an aggregate CE≈0 can be genuinely flat OR canceling rows (some repair, some break). Same forwards as the aggregate path — just skip the mean/median. - direct_total_effect(per_row=True): stores clean_rows/de_rows/te_rows (length n_test) - layerwise._gt_logit -> public gt_logit (used cross-module now) - run_direct_effect_sweep --per-row - plot_de_te_perrow.py: per-row hexbin + per-row CE histogram - test: per-row shape + aggregate DE == mean(per-row DE) on the linear coordinate Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
#44 Part 1 — result: no self-repair (path-patching DE vs TE · 4 models)What would show self-repair — 3 evidence, all required:
Logic: exists → mechanism → self-repair. Reference — what self-repair looks like (Hydra, McGrath 2023)E1 · below-diagonal downstream-repair mass (Δ_ablate vs Δ_unembed): E2 · tight compensation law — Layer 23, R²=0.92, slope=0.69: E1 — absent. DE–TE scatter (task×layer, margin, z-scored):
E1 robustness — is the null hidden by row-averaging? (per-row → no)Aggregate collapses ~500 rows/point → a null CE could hide a signal. The per-row CE distribution shape disambiguates the hypotheses:
E2 — absent. Hydra's claim = a systematic per-layer law (Fig-4d):
TFM analog — R² + slope vs depth, task-aggregated (15 independent pts/layer):
single-layer view — apex DE vs CE scatter (Hydra Fig-4b analog)Each model's peak-R² layer (final layer excluded — no downstream ⇒ CE≡0).
E3 — not reached, moot.
Takeaway:
|
5 PNGs referenced by raw.githubusercontent from the PR #45 comment: - fig1 aggregate DE–TE scatter, fig2 per-row CE hist, fig3 Hydra Fig-b regression - hydra_ref_scatter / hydra_ref_regression (McGrath 2023 reference) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
Task-aggregated (15 pts/layer, independent) — replaces the max-mean-CE single-layer strawman with the faithful per-layer R²+slope-vs-depth curve + apex (peak-R²) scatter. No layer reaches Hydra 0.92; the one structured layer (TabICL L10) is negative (breakage). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
| @@ -0,0 +1,121 @@ | |||
| """D1 — the DE–TE scatter (Hydra Fig 2c analog), from ``run_direct_effect_sweep.py``. | |||
There was a problem hiding this comment.
scripts/plot_de_te_scatter.py and scripts/plot_de_te_perrow.py?
这两个脚本什么关系 ?
| @@ -1 +1 @@ | |||
| """Run the exp6 self-repair sweep over the TabArena binary tasks. | |||
There was a problem hiding this comment.
name of this file should be updated, it is not self-repair as we discovered, let's find out a proper name
| @@ -1 +1 @@ | |||
| """Ablation sweep for the self-repair analysis. | |||
There was a problem hiding this comment.
is self_repair still correct? what does this script do?
Reproduce the published E1/E2 figures from code (figures kept in analysis/): - plot_ce_de_law.py — Hydra Fig-4d (per-layer CE~DE R²+slope vs depth) + Fig-4b apex scatter (--apex-scatter); --agg for task-aggregated (15 pts/layer). Peak-R² apex, final layer excluded (CE≡0). method-A CE is coupling-biased (noted in-file). - plot_de_te_perrow.py — add --hist-only (the per-row CE histogram, E1 robustness). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
…ing (audit follow-up) Minor items from a computation-logic audit (no bugs found): - layerwise_margin: note aggregate margin (median-of-medians) ≠ reduced per-row margin — median doesn't commute with subtraction; only gt_logit gives aggregate == mean(per-row). - plot_ce_de_law: docstring pointer to the same (margin --agg vs per-row are different estimators). - native_final_logits: comment the shared-decoder repeat (only [-1] read, stateless → safe). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
… fit (/simplify) Cleanup-review follow-up (behavior identical — all printed stats/figures unchanged): - new scripts/_de_te_common.py: MODEL_LABELS, load_de_json, de_scale — the de_<model>.json reader/labels/z-score rule were reimplemented in all 3 plot scripts. - plot_ce_de_law: _apex_layer recomputed _layer_curve that main already had → split into pure _apex_from_curve(r2, spread); no re-fit. _meaningful(spread) mask (was inlined 3×). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
…method-note comment The balef2026 self-repair observation: ablate a layer -> decoded AUC dips -> recovers to baseline by the final layer. Lead figure for the 'why dip-recover isn't self-repair' PR comment. (PNG only; the combined-plot script stays uncommitted.) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
Does dip→recover mean self-repair? (a note on balef2026's method)Ablate a layer, decode every depth with its tuned probe: the answer dips right after the ablated layer, then recovers to the no-ablation baseline by the final layer. Looks like the model repairs itself. Does it? — not from this measurement. The gap: What self-repair isDef (McGrath 2023, Hydra Effect): ablate a component → downstream actively compensates → net damage ≪ the component's own importance. Make "compensate" measurable — two counterfactuals for ablating layer m:
Why you need both — one number can't express it:
balef2026 measures only TE
Not that the conclusion is wrong — the method can't form the quantity that defines self-repair. Studying it requires the frozen-downstream DE. The paper stops where the question begins. Two further confounds in the tuned-decoder readEven granting a signal:
Measuring DE and TE on one ruler — and what it shows — is in the results comment. |
| @@ -0,0 +1,47 @@ | |||
| """Shared helpers for the DE/TE/CE plot scripts (``plot_de_te_scatter``, | |||
There was a problem hiding this comment.
is it a good pattern to name a script with prefix underscore? perhaps put those into tfm_lens utils or plot_utils, etc
…URLs) The 8 PNGs were only ever needed as a host for the #45 comment images. Those URLs are pinned to commits (40a6af6 / 9521d59 / dbc1629), which keep serving the blobs from history, so HEAD does not need to carry them. README.md records where each figure is pinned + how to regenerate it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
An underscore-prefixed module in scripts/ was a workaround for scripts/ not being a package. These helpers read the de_<model>.json contract that direct_effect.py writes, so they belong next to the producer — both ends of the contract now move together. No matplotlib in the module, so it is not "plot utils"; and utils.py is reserved for dependency-free helpers. All three figures re-generated: numbers unchanged. Addresses the review note on scripts/_de_te_common.py. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
…conds `pytest` ran every adapter integration test (real ckpt load + CPU forward) plus the finetune end-to-end path: minutes. Nothing distinguished them from the toy-adapter unit tests. - register the `real_model` marker; module-level on the 4 adapter suites and test_evaluation_real, test-level on the one real finetune path. - `pytest -m "not real_model"`: 91 passed in 3.5s (25 deselected). CI is unchanged (bare `pytest`): it still runs the LimiX-2M path for real (9 MB ckpt from HF); Mitra/TabICL/TabFM already skip there via importorskip. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
…46) (#48) * refactor: name things after the method, not the falsified conclusion (#46) #45 falsified self-repair in TFMs, so every `self_repair` name asserted a disproven hypothesis. Renames describe the mechanism; the paper-specific experiment now carries the paper's name. Library: - evaluation/self_repair.py -> evaluation/balef_exp6.py - evaluation/direct_effect.py -> evaluation/path_patching.py - direct_total_effect() -> layer_effects() - core/resample_ablation.py -> core/donor_delta.py (it only builds the delta) - new evaluation/native_readout.py: native_final_logits/auc extracted, since both path_patching and balef_exp6 use them — leaving them in the exp6 module would make its new name a lie. Scripts: - run_self_repair_sweep -> run_balef_exp6_sweep - plot_self_repair -> plot_balef_exp6_trajectory - plot_cross_dataset -> plot_balef_exp6_cross_dataset - run_direct_effect_sweep -> run_path_patching_sweep - plot_ce_de_law -> plot_compensation_fit (there is no law) - validate_resample_self_repair-> compare_ablation_stability - validate_resample_donor_health-> check_donor_delta_health Content, not just names: - compare_ablation_stability computed `SR = imm - TE` and called it self-repair. `imm` is a mid-depth tuned-probe reading, not the frozen-downstream DE, so that difference cannot identify self-repair. Renamed to `recovery` with the caveat stated; dropped the "Claim 1 — self-repair holds" framing. - neutralized docstrings/comments that asserted self-repair; kept the term where it names the hypothesis under test (e.g. the below-diagonal region). - evaluation/preprocess.py docstring claimed to be LimiX-only; it holds all four. Default artifact paths follow (out/self_repair*.json -> out/balef_exp6*.json). Tests renamed to match. 91 passed (-m "not real_model"). DESIGN.md / docs/reference-pipeline.md still mention experiments/exp6_self_repair.py — a historical layout that never existed in this tree; left alone. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa * docs: fix four leftovers found in review of the rename - README described Exp6 as quantifying downstream compensation. It does not — it measures per-depth decodability; compensation needs path_patching. - plot_balef_exp6_trajectory equated the dip-recover shape *with* self-repair, contradicting the module docstring it plots from. - compare_ablation_stability documented TE(m) and recovery(m), neither of which the script computes (only `imm` is). Dropped them and the caveat about a quantity that does not exist; the honest contract is imm + its spread/overshoot. - plot_compensation_fit still defaulted --out to out/ce_de_law.png. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The README still described the repo as a toolkit for reproducing Balef et al.'s Exp4/5/6. Since #45 the result is the other way round: the repo measures direct and total effects and finds no self-repair. - lead with the finding + the DE-TE scatter for LimiX-2M - state the DE/TE/CE definitions and the one-forward-pass shortcut - separate what reproduces from where the criterion differs - drop Exp5 (loop_layer is still a TODO -- it was never built) - update module and script names to their post-#48 form - link the write-up, the video, and the Hydra Effect paper Also record how the hero figure is generated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>







Part 1 of #44 — the DE–TE scatter (Hydra Fig 2c analog), plus a per-row extension. Additive: touches no existing module.
What & why
The scatter needs two rulers in one basis. This adds both, read through the model's own decoder (not per-layer tuned decoders) so they're comparable by construction.
DE(m) — direct effect, path-patching (method B). Residual is additive ⇒ freezing downstream at clean is arithmetic, not a per-layer forward:
One clean forward captures every
a_m; DE(m) for all m is residual arithmetic through the true decoder (final LN σ recomputed) — not a fixed û on one layer (that would be method A).TE(m) — total effect. Ablate-and-react (
inject_delta), read the final residual through the same decoder.DE and TE share one donor draw per (m, donor) → clean CE pairing for Part 2.
ã_m: resample — role-matched donor δ, metric averaged (reusing Ablation method: resample (cross-table) instead of zero/skip #35).Per-row (
--per-row): also store un-aggregated per-test-row DE/TE → closes the row-level sign-cancellation loophole (E1 robustness). Same forwards, just skip the mean/median.Files
src/tfm_lens/evaluation/direct_effect.py—direct_total_effect(+per_row) +final_decoder_logits(the shared ruler).src/tfm_lens/evaluation/layerwise.py—_gt_logit→ publicgt_logit(now used cross-module).scripts/run_direct_effect_sweep.py— per-task DE/TE + native-final normalizer →out/de_<model>.json(--per-rowoptional).scripts/plot_de_te_scatter.py— D1: DE(x) vs TE(y) + diagonal, per (layer, dataset), z-scored per task; redundantDE≈0stripe flagged.scripts/plot_de_te_perrow.py— per-row hexbin + CE histogram.tests/test_direct_effect.py.Tests
aggregate DE == mean(per-row DE)on the linear coordinate.einxabsent locally).Results
Full 4-model × 15-task run + analysis → no self-repair (falsification of #44). See the results comment.
Not in this PR (per #44)
🤖 Generated with Claude Code