refactor: name things after the method, not the falsified conclusion (#46) - #48
Merged
Merged
Conversation
…46) #45 falsified self-repair in TFMs, so every `self_repair` name asserted a disproven hypothesis. Renames describe the mechanism; the paper-specific experiment now carries the paper's name. Library: - evaluation/self_repair.py -> evaluation/balef_exp6.py - evaluation/direct_effect.py -> evaluation/path_patching.py - direct_total_effect() -> layer_effects() - core/resample_ablation.py -> core/donor_delta.py (it only builds the delta) - new evaluation/native_readout.py: native_final_logits/auc extracted, since both path_patching and balef_exp6 use them — leaving them in the exp6 module would make its new name a lie. Scripts: - run_self_repair_sweep -> run_balef_exp6_sweep - plot_self_repair -> plot_balef_exp6_trajectory - plot_cross_dataset -> plot_balef_exp6_cross_dataset - run_direct_effect_sweep -> run_path_patching_sweep - plot_ce_de_law -> plot_compensation_fit (there is no law) - validate_resample_self_repair-> compare_ablation_stability - validate_resample_donor_health-> check_donor_delta_health Content, not just names: - compare_ablation_stability computed `SR = imm - TE` and called it self-repair. `imm` is a mid-depth tuned-probe reading, not the frozen-downstream DE, so that difference cannot identify self-repair. Renamed to `recovery` with the caveat stated; dropped the "Claim 1 — self-repair holds" framing. - neutralized docstrings/comments that asserted self-repair; kept the term where it names the hypothesis under test (e.g. the below-diagonal region). - evaluation/preprocess.py docstring claimed to be LimiX-only; it holds all four. Default artifact paths follow (out/self_repair*.json -> out/balef_exp6*.json). Tests renamed to match. 91 passed (-m "not real_model"). DESIGN.md / docs/reference-pipeline.md still mention experiments/exp6_self_repair.py — a historical layout that never existed in this tree; left alone. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
- README described Exp6 as quantifying downstream compensation. It does not — it measures per-depth decodability; compensation needs path_patching. - plot_balef_exp6_trajectory equated the dip-recover shape *with* self-repair, contradicting the module docstring it plots from. - compare_ablation_stability documented TE(m) and recovery(m), neither of which the script computes (only `imm` is). Dropped them and the caveat about a quantity that does not exist; the honest contract is imm + its spread/overshoot. - plot_compensation_fit still defaulted --out to out/ce_de_law.png. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L4s6SJPSt2vQrw1pE7wBxa
xiaohan2012
added a commit
that referenced
this pull request
Sep 3, 2026
The README still described the repo as a toolkit for reproducing Balef et al.'s Exp4/5/6. Since #45 the result is the other way round: the repo measures direct and total effects and finds no self-repair. - lead with the finding + the DE-TE scatter for LimiX-2M - state the DE/TE/CE definitions and the one-forward-pass shortcut - separate what reproduces from where the criterion differs - drop Exp5 (loop_layer is still a TODO -- it was never built) - update module and script names to their post-#48 form - link the write-up, the video, and the Hydra Effect paper Also record how the hero figure is generated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #46.
#45 falsified self-repair in TFMs, so every
self_repairname asserted a disproven hypothesis. Rule applied throughout: name the method/mechanism, never the conclusion; a paper-specific experiment carries the paper's name.Library
evaluation/self_repair.pyevaluation/balef_exp6.pyevaluation/direct_effect.pyevaluation/path_patching.pydirect_total_effect()layer_effects()path_patching.layer_effects(...)core/resample_ablation.pycore/donor_delta.pycore/interventions.pyexecutes the ablationevaluation/native_readout.pynative_final_logits/aucextractedThe extraction is a prerequisite, not a bonus: both
path_patchingandbalef_exp6use those two functions. Leaving them inside the exp6 module would make its new name a lie.Scripts
run_self_repair_sweep.pyrun_balef_exp6_sweep.pyplot_self_repair.pyplot_balef_exp6_trajectory.pyplot_cross_dataset.pyplot_balef_exp6_cross_dataset.pyrun_direct_effect_sweep.pyrun_path_patching_sweep.pyplot_ce_de_law.pyplot_compensation_fit.py— our result is that there is no lawvalidate_resample_self_repair.pycompare_ablation_stability.pyvalidate_resample_donor_health.pycheck_donor_delta_health.pycompare_ablation_stabilityevaluates the perturbation method (std(imm)across layers + overshoot count), not Exp6's conclusion — and it explicitly refuses to treat zero as ground truth, so "validate … vs …" was the wrong verb.Content, not just filenames
compare_ablation_stabilitycomputedSR(m) = imm(m) − TE(m)and called it self-repair. That is exactly the non-identifiable quantity:immis a mid-depth tuned-probe reading, not the frozen-downstream DE. Renamed torecovery, caveat stated inline,"Claim 1 — self-repair holds"framing dropped.evaluation/preprocess.pydocstring claimed LimiX-only; the module holds all four backbones.out/self_repair*.json→out/balef_exp6*.json.Tests
test_self_repair→test_balef_exp6·test_direct_effect→test_path_patching·test_resample_ablation→test_donor_delta.91 passed, 1 skipped (
-m "not real_model", 3.5s). No behavior change — renames + docstrings only.Left alone
DESIGN.md/docs/reference-pipeline.mdreferenceexperiments/exp6_self_repair.py, a historical layout that never existed in this tree.plot_hydra_layer.py,plot_ce_vs_de.py,plot_layer_ce.py,draw_ce_methods.py) are untracked — author's call to delete.🤖 Generated with Claude Code