Skip to content

feat: path-patching DE/TE on the native head — DE–TE scatter (#44, Part 1) - #45

Merged
xiaohan2012 merged 13 commits into
mainfrom
feat/direct-effect
Aug 5, 2026
Merged

xiaohan2012 merged 13 commits into
mainfrom
feat/direct-effect

Conversation

@xiaohan2012

@xiaohan2012 xiaohan2012 commented Aug 4, 2026 •

Copy link
Copy Markdown
Owner

Part 1 of #44 — the DE–TE scatter (Hydra Fig 2c analog), plus a per-row extension. Additive: touches no existing module.

What & why

The scatter needs two rulers in one basis. This adds both, read through the model's own decoder (not per-layer tuned decoders) so they're comparable by construction.

  • DE(m) — direct effect, path-patching (method B). Residual is additive ⇒ freezing downstream at clean is arithmetic, not a per-layer forward:

    r_L^DE(m) = r_L^clean − a_m + ã_m
    

    One clean forward captures every a_m; DE(m) for all m is residual arithmetic through the true decoder (final LN σ recomputed) — not a fixed û on one layer (that would be method A).

  • TE(m) — total effect. Ablate-and-react (inject_delta), read the final residual through the same decoder.

  • DE and TE share one donor draw per (m, donor) → clean CE pairing for Part 2.

  • ã_m: resample — role-matched donor δ, metric averaged (reusing Ablation method: resample (cross-table) instead of zero/skip #35).

  • Per-row (--per-row): also store un-aggregated per-test-row DE/TE → closes the row-level sign-cancellation loophole (E1 robustness). Same forwards, just skip the mean/median.

Files

  • src/tfm_lens/evaluation/direct_effect.py — direct_total_effect (+ per_row) + final_decoder_logits (the shared ruler).
  • src/tfm_lens/evaluation/layerwise.py — _gt_logit → public gt_logit (now used cross-module).
  • scripts/run_direct_effect_sweep.py — per-task DE/TE + native-final normalizer → out/de_<model>.json (--per-row optional).
  • scripts/plot_de_te_scatter.py — D1: DE(x) vs TE(y) + diagonal, per (layer, dataset), z-scored per task; redundant DE≈0 stripe flagged.
  • scripts/plot_de_te_perrow.py — per-row hexbin + CE histogram.
  • tests/test_direct_effect.py.

Tests

  • Headline invariant: the O(1)-forward arithmetic DE == a real frozen-downstream forward (skip m, pin every downstream layer to its clean contribution). Parametrized over 3D / 4D / double-stream — if these match, the shortcut is exact for any decoder/layout.
  • DE == logit-lens closed form on a linear head (method A == B).
  • Per-row: shape + aggregate DE == mean(per-row DE) on the linear coordinate.
  • resample path shape/keys/finiteness; missing-donor guard.
  • Full suite: 109 passed, 8 skipped (skips pre-existing: einx absent locally).

Results

Full 4-model × 15-task run + analysis → no self-repair (falsification of #44). See the results comment.

Not in this PR (per #44)

  • Part 2: CE decomposition + CE~DE regression (independent CE).
  • Part 3: LN-artifact (RQ3) fold-σ subtraction.

🤖 Generated with Claude Code

…t 1)

The DE–TE scatter (Hydra Fig 2c analog) needs two rulers in one basis. Add
`evaluation/direct_effect.py`: DE (direct effect, path-patching 取法 B) and TE
(total effect), both read through the model's own head — not per-layer tuned
decoders — so they're comparable by construction.

- DE(m): residual is additive, so freezing downstream at clean is arithmetic,
  r_L − a_m + ã_m, off one clean forward. Faithful head (final LN σ recomputed),
  not a fixed û on a single layer.
- TE(m): ablate-and-react (inject_delta / skip_layer), read the final residual
  through the same head. DE and TE share one donor draw per (m, donor).
- resample ã_m (role-matched donor δ, metric averaged) reusing the #35 machinery;
  zero as the cross-check.

Plus `run_direct_effect_sweep.py` (per-task DE/TE + native-final normalizer) and
`plot_de_te_scatter.py` (D1). Additive PR: touches no existing module.

Tests (tests/test_direct_effect.py): the headline invariant is that the O(1)-forward
arithmetic DE reproduces a real frozen-downstream forward (3D / 4D / double-stream);
plus DE == logit-lens on a linear head (取法 A == B) and the resample path's shape.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa

@xiaohan2012 xiaohan2012 left a comment •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left some comments.

Other remarks:

  • add type hints to improve readability.
  • perhaps move some common util functions to some utility script

Comment thread src/tfm_lens/evaluation/direct_effect.py Outdated
Comment thread src/tfm_lens/evaluation/direct_effect.py Outdated
Comment thread src/tfm_lens/evaluation/direct_effect.py Outdated
Comment thread src/tfm_lens/evaluation/direct_effect.py Outdated
Comment thread src/tfm_lens/evaluation/direct_effect.py Outdated
Comment thread src/tfm_lens/evaluation/direct_effect.py Outdated
Comment thread src/tfm_lens/evaluation/direct_effect.py Outdated
Comment thread src/tfm_lens/evaluation/direct_effect.py Outdated
- rename for readability: native_head_logits -> final_decoder_logits, _swap ->
  patched_residual; descriptive vars (clean_contributions / clean_residual /
  replacement); unify "head" -> "decoder"; document a_m = clean write r_m - r_{m-1}.
- drop zero ablation entirely (module + sweep + tests) — resample only, as agreed.
- dedup: move clone_residual to utils (was self_repair._clone_residual); factor
  native_final_logits out of native_final_auc so TE reuses it instead of a bespoke
  forward (_te_forward deleted).
- add type hints (Residual / Coords aliases, signatures + returns).

Tests rewritten to the primitive-level invariant with an arbitrary replacement ã_m
(no zero): patched_residual == frozen-downstream forward (3D / 4D / double-stream),
plus the linear-head 取法 A == B identity. Re-smoked on LimiX-2M end-to-end.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
@xiaohan2012

xiaohan2012 commented Aug 4, 2026 •

Copy link
Copy Markdown
Owner Author

Review pass in df5077f. Summary:

  • Renames / clarity: native_head_logits→final_decoder_logits, _swap→patched_residual, descriptive vars (clean_contributions / clean_residual / replacement), "head"→"decoder" throughout. a_m documented = clean write r_m − r_{m-1}.
  • No zero ablation: removed from module, sweep, and tests — resample only.
  • Dedup / utils: clone_residual → utils.py (was self_repair._clone_residual); native_final_logits factored out of native_final_auc so TE reuses it (deleted the bespoke _te_forward).
  • Type hints: added everywhere, incl. Residual / Coords aliases.

Tests rewritten to the primitive invariant with an arbitrary ã_m (no zero): patched_residual == frozen-downstream forward (3D/4D/double-stream) + linear-head A==B. Full suite green; re-smoked on LimiX-2M end-to-end.

Comment thread scripts/run_direct_effect_sweep.py Outdated
Comment thread src/tfm_lens/evaluation/direct_effect.py Outdated
Comment thread src/tfm_lens/evaluation/direct_effect.py Outdated
Comment thread src/tfm_lens/evaluation/direct_effect.py Outdated
xiaohan2012 and others added 2 commits August 4, 2026 15:32
- naming: Coords -> MetricPair, _coords -> _metric_pair, _effect -> _drop_from_clean,
  _mean_coords -> _mean_over_donors (more specific names per review).
- English: translate the remaining 取法 A/B -> method A/B across module, scripts, tests.
- dedup: move the duplicated _subsample / _load_record out of both sweep scripts into
  evaluation/datasets.py as subsample() + load_task_record(); scripts import them.

Re-smoked on LimiX-2M end-to-end; ruff + affected tests green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
Un-aggregate the ~500 test rows per (task, layer) to close the sign-cancellation
loophole: an aggregate CE≈0 can be genuinely flat OR canceling rows (some repair,
some break). Same forwards as the aggregate path — just skip the mean/median.

- direct_total_effect(per_row=True): stores clean_rows/de_rows/te_rows (length n_test)
- layerwise._gt_logit -> public gt_logit (used cross-module now)
- run_direct_effect_sweep --per-row
- plot_de_te_perrow.py: per-row hexbin + per-row CE histogram
- test: per-row shape + aggregate DE == mean(per-row DE) on the linear coordinate

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
@xiaohan2012

xiaohan2012 commented Aug 5, 2026 •

Copy link
Copy Markdown
Owner Author

#44 Part 1 — result: no self-repair (path-patching DE vs TE · 4 models)

What would show self-repair — 3 evidence, all required:

evidence signature
E1 decoupling at real DE below-diagonal cluster (TE≪DE) with large DE (not the DE≈0 stripe)
E2 systematic law tight CE~DE line (Hydra R² 0.92, slope 0.69)
E3 active, not artifact localizable to layers + survives LN-σ subtraction

Logic: exists → mechanism → self-repair.

Reference — what self-repair looks like (Hydra, McGrath 2023)

E1 · below-diagonal downstream-repair mass (Δ_ablate vs Δ_unembed):

Hydra Δ_ablate vs Δ_unembed — downstream-repair mass

E2 · tight compensation law — Layer 23, R²=0.92, slope=0.69:

Hydra Layer 23 — R²=0.92, slope=0.69


E1 — absent. DE–TE scatter (task×layer, margin, z-scored):

DE–TE scatter (task×layer, margin, z-scored)

  • 70–90% of points in the DE≈0 redundant stripe (non-redundant: LimiX 55/180, Mitra 28/180, TabICL 34/180, TabFM 39/360) → no large-DE to decouple.
  • rest: mean CE ≈ 0σ (+0.09 / +0.01 / −0.08 / +0.00); below-diag% noisy 35–68%, no direction.
E1 robustness — is the null hidden by row-averaging? (per-row → no)

Aggregate collapses ~500 rows/point → a null CE could hide a signal. The per-row CE distribution shape disambiguates the hypotheses:

per-row CE shape hypothesis self-repair?
single peak at 0 genuinely flat — no compensation no (honest null)
bimodal / two fat tails sign cancellation (½ CE>0 repair, ½ CE<0 break) would hide repair
right-skew / fat +tail diluted subpopulation (repair on a few hard rows) would hide repair
+mean, unimodal uniform repair yes
left-skew / −mean net breakage / amplification opposite of repair

per-row CE histogram (margin) — unimodal at 0

  • observed: single peak at 0 for all 4 → genuinely flat, not cancellation, not a diluted subpopulation.
  • per-task CE / 15 tables: LimiX +0.06±0.14 (12/15 pos, but 8/15 on gt_logit → not coord-stable), Mitra −0.12 (3/15), TabICL −0.23 (1/15), TabFM +0.00 (10/15).
  • only stable signal = TabICL CE<0 (left-lean, amplification), opposite of repair.

E2 — absent. Hydra's claim = a systematic per-layer law (Fig-4d):

  • CE~DE regression R² + slope form a band across middle-late layers.
  • apex at L23: R²=0.92, slope=0.69.

TFM analog — R² + slope vs depth, task-aggregated (15 independent pts/layer):

per-layer R² + slope vs depth, task-aggregated (margin)

model apex (peak R²) slope R²
LimiX-2M L10 −0.03 0.00
Mitra L10 +0.08 0.04
TabICLv2 L10 −0.41 0.58 (negative = amplification)
TabFM L18 −0.08 0.04
  • no layer has both R²→0.92 and slope∈(0,1). Most layers are redundant (DE≈0, hollow markers).
  • the one structured layer (TabICL L10, R²0.58) has a negative slope = downstream amplification, opposite of repair.
  • per-row regression (Hydra's per-prompt analog, thousands of pts) agrees — apex layers negative/weak either way.
  • CE=DE−TE is method-A (mechanical coupling → biased toward a positive law); even so, no Hydra band. Coupling-free version = independent CE (Part 2).
single-layer view — apex DE vs CE scatter (Hydra Fig-4b analog)

apex-layer DE vs CE scatter + fit, task-aggregated (margin)

Each model's peak-R² layer (final layer excluded — no downstream ⇒ CE≡0).

  • Hydra L23 = tight positive cloud below the 1:1 line.
  • ours = flat blobs (R²≈0), or TabICL's downward line (amplification).

E3 — not reached, moot.

  • E1/E2 absent → nothing to attribute/artifact-correct.
  • LN branch pre-empted by Mitra (linear head, ε=0, also null) → the null isn't LN masking a real effect.
  • per-layer CE + σ-freeze = remaining RQ3/Part-2; only if a signal appears.

Takeaway:

xiaohan2012 and others added 2 commits August 5, 2026 16:06
5 PNGs referenced by raw.githubusercontent from the PR #45 comment:
- fig1 aggregate DE–TE scatter, fig2 per-row CE hist, fig3 Hydra Fig-b regression
- hydra_ref_scatter / hydra_ref_regression (McGrath 2023 reference)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
Task-aggregated (15 pts/layer, independent) — replaces the max-mean-CE single-layer
strawman with the faithful per-layer R²+slope-vs-depth curve + apex (peak-R²) scatter.
No layer reaches Hydra 0.92; the one structured layer (TabICL L10) is negative (breakage).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
@@ -0,0 +1,121 @@
"""D1 — the DE–TE scatter (Hydra Fig 2c analog), from ``run_direct_effect_sweep.py``.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

scripts/plot_de_te_scatter.py and scripts/plot_de_te_perrow.py?

这两个脚本什么关系 ?

@@ -1 +1 @@
"""Run the exp6 self-repair sweep over the TabArena binary tasks.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

name of this file should be updated, it is not self-repair as we discovered, let's find out a proper name

@@ -1 +1 @@
"""Ablation sweep for the self-repair analysis.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is self_repair still correct? what does this script do?

xiaohan2012 and others added 4 commits August 5, 2026 17:34
Reproduce the published E1/E2 figures from code (figures kept in analysis/):
- plot_ce_de_law.py — Hydra Fig-4d (per-layer CE~DE R²+slope vs depth) + Fig-4b apex
  scatter (--apex-scatter); --agg for task-aggregated (15 pts/layer). Peak-R² apex,
  final layer excluded (CE≡0). method-A CE is coupling-biased (noted in-file).
- plot_de_te_perrow.py — add --hist-only (the per-row CE histogram, E1 robustness).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
…ing (audit follow-up)

Minor items from a computation-logic audit (no bugs found):
- layerwise_margin: note aggregate margin (median-of-medians) ≠ reduced per-row margin —
  median doesn't commute with subtraction; only gt_logit gives aggregate == mean(per-row).
- plot_ce_de_law: docstring pointer to the same (margin --agg vs per-row are different estimators).
- native_final_logits: comment the shared-decoder repeat (only [-1] read, stateless → safe).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
… fit (/simplify)

Cleanup-review follow-up (behavior identical — all printed stats/figures unchanged):
- new scripts/_de_te_common.py: MODEL_LABELS, load_de_json, de_scale — the de_<model>.json
  reader/labels/z-score rule were reimplemented in all 3 plot scripts.
- plot_ce_de_law: _apex_layer recomputed _layer_curve that main already had → split into
  pure _apex_from_curve(r2, spread); no re-fit. _meaningful(spread) mask (was inlined 3×).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
…method-note comment

The balef2026 self-repair observation: ablate a layer -> decoded AUC dips -> recovers
to baseline by the final layer. Lead figure for the 'why dip-recover isn't self-repair'
PR comment. (PNG only; the combined-plot script stays uncommitted.)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
@xiaohan2012

Copy link
Copy Markdown
Owner Author

Does dip→recover mean self-repair? (a note on balef2026's method)

Ablate a layer, decode every depth with its tuned probe: the answer dips right after the ablated layer, then recovers to the no-ablation baseline by the final layer.

dip→recover across 4 TFMs (exp6, tuned decoders)

Looks like the model repairs itself. Does it? — not from this measurement. The gap:

What self-repair is

Def (McGrath 2023, Hydra Effect): ablate a component → downstream actively compensates → net damage ≪ the component's own importance.

Make "compensate" measurable — two counterfactuals for ablating layer m:

  • DE (direct effect) — remove m's write, freeze downstream at clean. What m directly contributes (damage if nobody reacts).
  • TE (total effect) — remove m, let downstream react. The net damage.
  • CE = DE − TE — what downstream made up. self-repair ⟺ CE > 0 (TE ≪ DE).

Why you need both — one number can't express it:

  • TE alone is ambiguous. TE≈0 = layer never mattered (redundant) or mattered and was fully repaired. Same reading, opposite mechanisms.
  • DE alone is ambiguous. DE>0 = load-bearing (TE=DE) or repaired (TE≪DE).
  • Only the gap DE − TE isolates compensation — "downstream compensated" is "direct importance > net importance", a two-quantity comparison.

Mediation view: DE = direct path m→out, TE = total, CE = indirect path through downstream. Self-repair = the mediator offsetting the direct damage.

balef2026 measures only TE

  • dip→recover is a per-depth decoded trajectory. Its "recover to baseline" is TE≈0 — net damage at the final layer, downstream already free to react.
  • It never freezes downstream → never produces DE.
  • So it reads exactly TE≈0 — the ambiguous case — as repair. But TE≈0 is equally redundancy; only DE separates them.
DE TE verdict
balef's "recover" not measured ≈0 redundant or repaired — undecidable

Not that the conclusion is wrong — the method can't form the quantity that defines self-repair. Studying it requires the frozen-downstream DE. The paper stops where the question begins.

Two further confounds in the tuned-decoder read

Even granting a signal:

  1. Decodability ≠ causal contribution. A per-depth tuned probe measures how well the answer can be decoded from the residual — info present, not info used by the native readout. And each depth uses a different probe → dip and recover are read on different rulers.
  2. dip→recover ⊂ passive redundancy. If the answer is already duplicated in the residual, deleting one copy dips a local probe but the others survive → the final readout still decodes it. Nobody rewrote anything; later layers pick up redundant info. (Cheap in TFMs — the input table stays in context.)

Measuring DE and TE on one ruler — and what it shows — is in the results comment.

Comment thread scripts/_de_te_common.py Outdated
@@ -0,0 +1,47 @@
"""Shared helpers for the DE/TE/CE plot scripts (``plot_de_te_scatter``,

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is it a good pattern to name a script with prefix underscore? perhaps put those into tfm_lens utils or plot_utils, etc

xiaohan2012 and others added 3 commits August 5, 2026 21:33
…URLs)

The 8 PNGs were only ever needed as a host for the #45 comment images.
Those URLs are pinned to commits (40a6af6 / 9521d59 / dbc1629), which keep
serving the blobs from history, so HEAD does not need to carry them.

README.md records where each figure is pinned + how to regenerate it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
An underscore-prefixed module in scripts/ was a workaround for scripts/ not
being a package. These helpers read the de_<model>.json contract that
direct_effect.py writes, so they belong next to the producer — both ends of
the contract now move together. No matplotlib in the module, so it is not
"plot utils"; and utils.py is reserved for dependency-free helpers.

All three figures re-generated: numbers unchanged.

Addresses the review note on scripts/_de_te_common.py.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
…conds

`pytest` ran every adapter integration test (real ckpt load + CPU forward)
plus the finetune end-to-end path: minutes. Nothing distinguished them from
the toy-adapter unit tests.

- register the `real_model` marker; module-level on the 4 adapter suites and
  test_evaluation_real, test-level on the one real finetune path.
- `pytest -m "not real_model"`: 91 passed in 3.5s (25 deselected).

CI is unchanged (bare `pytest`): it still runs the LimiX-2M path for real
(9 MB ckpt from HF); Mitra/TabICL/TabFM already skip there via importorskip.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
@xiaohan2012
xiaohan2012 merged commit 5fa8b34 into main Aug 5, 2026
1 check passed
@xiaohan2012
xiaohan2012 deleted the feat/direct-effect branch August 5, 2026 18:43
xiaohan2012 added a commit that referenced this pull request Aug 5, 2026
…46) (#48)

* refactor: name things after the method, not the falsified conclusion (#46)

#45 falsified self-repair in TFMs, so every `self_repair` name asserted a
disproven hypothesis. Renames describe the mechanism; the paper-specific
experiment now carries the paper's name.

Library:
- evaluation/self_repair.py    -> evaluation/balef_exp6.py
- evaluation/direct_effect.py  -> evaluation/path_patching.py
- direct_total_effect()        -> layer_effects()
- core/resample_ablation.py    -> core/donor_delta.py  (it only builds the delta)
- new evaluation/native_readout.py: native_final_logits/auc extracted, since
  both path_patching and balef_exp6 use them — leaving them in the exp6 module
  would make its new name a lie.

Scripts:
- run_self_repair_sweep        -> run_balef_exp6_sweep
- plot_self_repair             -> plot_balef_exp6_trajectory
- plot_cross_dataset           -> plot_balef_exp6_cross_dataset
- run_direct_effect_sweep      -> run_path_patching_sweep
- plot_ce_de_law               -> plot_compensation_fit  (there is no law)
- validate_resample_self_repair-> compare_ablation_stability
- validate_resample_donor_health-> check_donor_delta_health

Content, not just names:
- compare_ablation_stability computed `SR = imm - TE` and called it self-repair.
  `imm` is a mid-depth tuned-probe reading, not the frozen-downstream DE, so that
  difference cannot identify self-repair. Renamed to `recovery` with the caveat
  stated; dropped the "Claim 1 — self-repair holds" framing.
- neutralized docstrings/comments that asserted self-repair; kept the term where
  it names the hypothesis under test (e.g. the below-diagonal region).
- evaluation/preprocess.py docstring claimed to be LimiX-only; it holds all four.

Default artifact paths follow (out/self_repair*.json -> out/balef_exp6*.json).
Tests renamed to match. 91 passed (-m "not real_model").

DESIGN.md / docs/reference-pipeline.md still mention experiments/exp6_self_repair.py
— a historical layout that never existed in this tree; left alone.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa

* docs: fix four leftovers found in review of the rename

- README described Exp6 as quantifying downstream compensation. It does not —
  it measures per-depth decodability; compensation needs path_patching.
- plot_balef_exp6_trajectory equated the dip-recover shape *with* self-repair,
  contradicting the module docstring it plots from.
- compare_ablation_stability documented TE(m) and recovery(m), neither of which
  the script computes (only `imm` is). Dropped them and the caveat about a
  quantity that does not exist; the honest contract is imm + its spread/overshoot.
- plot_compensation_fit still defaulted --out to out/ce_de_law.png.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
xiaohan2012 added a commit that referenced this pull request Sep 3, 2026
The README still described the repo as a toolkit for reproducing Balef et
al.'s Exp4/5/6. Since #45 the result is the other way round: the repo
measures direct and total effects and finds no self-repair.

- lead with the finding + the DE-TE scatter for LimiX-2M
- state the DE/TE/CE definitions and the one-forward-pass shortcut
- separate what reproduces from where the criterion differs
- drop Exp5 (loop_layer is still a TODO -- it was never built)
- update module and script names to their post-#48 form
- link the write-up, the video, and the Hydra Effect paper

Also record how the hero figure is generated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant