Skip to content

refactor: name things after the method, not the falsified conclusion (#46) - #48

Merged
xiaohan2012 merged 2 commits into
mainfrom
refactor/faithful-naming
Aug 5, 2026
Merged

xiaohan2012 merged 2 commits into
mainfrom
refactor/faithful-naming

Conversation

@xiaohan2012

Copy link
Copy Markdown
Owner

Closes #46.

#45 falsified self-repair in TFMs, so every self_repair name asserted a disproven hypothesis. Rule applied throughout: name the method/mechanism, never the conclusion; a paper-specific experiment carries the paper's name.

Library

from to why
evaluation/self_repair.py evaluation/balef_exp6.py the per-depth-tuned-decoder design is balef2026's, not a generic ablation
evaluation/direct_effect.py evaluation/path_patching.py the module's identity is the method; it returns both DE and TE — old name dropped half
direct_total_effect() layer_effects() reads as path_patching.layer_effects(...)
core/resample_ablation.py core/donor_delta.py it only builds δ; core/interventions.py executes the ablation
— new evaluation/native_readout.py native_final_logits/auc extracted

The extraction is a prerequisite, not a bonus: both path_patching and balef_exp6 use those two functions. Leaving them inside the exp6 module would make its new name a lie.

Scripts

from to
run_self_repair_sweep.py run_balef_exp6_sweep.py
plot_self_repair.py plot_balef_exp6_trajectory.py
plot_cross_dataset.py plot_balef_exp6_cross_dataset.py
run_direct_effect_sweep.py run_path_patching_sweep.py
plot_ce_de_law.py plot_compensation_fit.py — our result is that there is no law
validate_resample_self_repair.py compare_ablation_stability.py
validate_resample_donor_health.py check_donor_delta_health.py

compare_ablation_stability evaluates the perturbation method (std(imm) across layers + overshoot count), not Exp6's conclusion — and it explicitly refuses to treat zero as ground truth, so "validate … vs …" was the wrong verb.

Content, not just filenames

  • compare_ablation_stability computed SR(m) = imm(m) − TE(m) and called it self-repair. That is exactly the non-identifiable quantity: imm is a mid-depth tuned-probe reading, not the frozen-downstream DE. Renamed to recovery, caveat stated inline, "Claim 1 — self-repair holds" framing dropped.
  • Neutralized docstrings/comments that asserted self-repair. Kept the term where it names the hypothesis under test (e.g. "below-diagonal = self-repair" in the DE–TE scatter) — that usage is correct.
  • evaluation/preprocess.py docstring claimed LimiX-only; the module holds all four backbones.
  • Default artifact paths follow: out/self_repair*.json → out/balef_exp6*.json.

Tests

test_self_repair → test_balef_exp6 · test_direct_effect → test_path_patching · test_resample_ablation → test_donor_delta.

91 passed, 1 skipped (-m "not real_model", 3.5s). No behavior change — renames + docstrings only.

Left alone

  • DESIGN.md / docs/reference-pipeline.md reference experiments/exp6_self_repair.py, a historical layout that never existed in this tree.
  • Superseded local scripts (plot_hydra_layer.py, plot_ce_vs_de.py, plot_layer_ce.py, draw_ce_methods.py) are untracked — author's call to delete.

🤖 Generated with Claude Code

xiaohan2012 and others added 2 commits August 5, 2026 21:59
…46)

#45 falsified self-repair in TFMs, so every `self_repair` name asserted a
disproven hypothesis. Renames describe the mechanism; the paper-specific
experiment now carries the paper's name.

Library:
- evaluation/self_repair.py    -> evaluation/balef_exp6.py
- evaluation/direct_effect.py  -> evaluation/path_patching.py
- direct_total_effect()        -> layer_effects()
- core/resample_ablation.py    -> core/donor_delta.py  (it only builds the delta)
- new evaluation/native_readout.py: native_final_logits/auc extracted, since
  both path_patching and balef_exp6 use them — leaving them in the exp6 module
  would make its new name a lie.

Scripts:
- run_self_repair_sweep        -> run_balef_exp6_sweep
- plot_self_repair             -> plot_balef_exp6_trajectory
- plot_cross_dataset           -> plot_balef_exp6_cross_dataset
- run_direct_effect_sweep      -> run_path_patching_sweep
- plot_ce_de_law               -> plot_compensation_fit  (there is no law)
- validate_resample_self_repair-> compare_ablation_stability
- validate_resample_donor_health-> check_donor_delta_health

Content, not just names:
- compare_ablation_stability computed `SR = imm - TE` and called it self-repair.
  `imm` is a mid-depth tuned-probe reading, not the frozen-downstream DE, so that
  difference cannot identify self-repair. Renamed to `recovery` with the caveat
  stated; dropped the "Claim 1 — self-repair holds" framing.
- neutralized docstrings/comments that asserted self-repair; kept the term where
  it names the hypothesis under test (e.g. the below-diagonal region).
- evaluation/preprocess.py docstring claimed to be LimiX-only; it holds all four.

Default artifact paths follow (out/self_repair*.json -> out/balef_exp6*.json).
Tests renamed to match. 91 passed (-m "not real_model").

DESIGN.md / docs/reference-pipeline.md still mention experiments/exp6_self_repair.py
— a historical layout that never existed in this tree; left alone.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
- README described Exp6 as quantifying downstream compensation. It does not —
  it measures per-depth decodability; compensation needs path_patching.
- plot_balef_exp6_trajectory equated the dip-recover shape *with* self-repair,
  contradicting the module docstring it plots from.
- compare_ablation_stability documented TE(m) and recovery(m), neither of which
  the script computes (only `imm` is). Dropped them and the caveat about a
  quantity that does not exist; the honest contract is imm + its spread/overshoot.
- plot_compensation_fit still defaulted --out to out/ce_de_law.png.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
@xiaohan2012
xiaohan2012 merged commit cedf6a3 into main Aug 5, 2026
1 check passed
@xiaohan2012
xiaohan2012 deleted the refactor/faithful-naming branch August 5, 2026 19:07
xiaohan2012 added a commit that referenced this pull request Sep 3, 2026
The README still described the repo as a toolkit for reproducing Balef et
al.'s Exp4/5/6. Since #45 the result is the other way round: the repo
measures direct and total effects and finds no self-repair.

- lead with the finding + the DE-TE scatter for LimiX-2M
- state the DE/TE/CE definitions and the one-forward-pass shortcut
- separate what reproduces from where the criterion differs
- drop Exp5 (loop_layer is still a TODO -- it was never built)
- update module and script names to their post-#48 form
- link the write-up, the video, and the Hydra Effect paper

Also record how the hero figure is generated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Faithful naming: describe the method, not the falsified conclusion

1 participant