Working thesis: v2 showed graph and lexical signals are complementary but naive mixing degrades and rank fusion loses; v3 explains why (scale incommensurability), cures it by calibrating signals into one probabilistic scale, generalises EGO/PPR as two points of one diffusion family (HKPR), and replaces the heuristic diversity bonus with DPP selection.
Order of work
Prerequisites v3 prerequisites: fixes that distort the v3 measurements (§6) #245 (items 2, 5, 6 still open; --budget is documented as a hard cap three times and is not one: cores bypass it (3.15x at --budget 1000) #241 , Per-file budget share: one generated data file takes 88% of a 15-file range #238 alongside) → bitcheck / corpus → freeze
Data & protocol v3 data & protocol: dcbench calib/holdout split, pre-registration P1, tag paper-v3 #252 : dcbench calib/holdout split, frozen manifests, pre-registration, tag paper-v3 (The release-cycle equivalence gate is unreproducible off one laptop; 1150 lines of eval CI have never run #233 is a blocker here)
Telemetry v3 telemetry: per-file calibration features + coverage fields in harness output (E-class) #253 (E-class) — per-file features + coverage fields
C1 calibration v3 C1: per-mode isotonic calibration + calibrated fusion (fused-sum / noisy-OR) #246 → cells 9–12
C2 HKPR v3 C2: Heat Kernel PageRank (hk-relax) as the EGO/PPR diffusion family #247 (parallel with 4) → cells 4–8
C3 DPP v3 C3: DPP-MAP re-ranking replaces the path-dependent relatedness bonus #248 over the winner of 11/12
C5 discovery-union v3 C5: discovery-union + universe/selection coverage decomposition #250 + coverage decomposition (decides how much of Multi-hop import closure (depth 2-3): raise the universe ceiling #130 v3 needs); C4 sequential v3 C4: budget-compliant sequential fusion BM25→EGO #249 ; seeds v3: restore change_magnitude-weighted seeds (lost v1 signal) #251
Matrix v3 experiment matrix: 16 cells, budget curve, per-role table #254 — once, on holdout + public, per pre-registration
Paper v3 paper: skeleton, risks write-up, artifacts #255
E/Q discipline
All C1–C5 are Q-class → one eval cycle: prerequisites first, freeze, then the matrix. A Q-class change landing after the freeze invalidates calibration and forces a full rerun.
Tracker changes made 2026-08-30
Superseded by v3 C1: per-mode isotonic calibration + calibrated fusion (fused-sum / noisy-OR) #246 : Learned fusion (LTR) over per-candidate features — make-or-break #129 (LTR), Per-mode (tau, beta_core) re-tuning sensitivity grid #140 (per-mode τ/β re-tuning).
Closed as done: Harness: whitespace-tolerant apply fallback (recover 5 PolyBench CRLF rows) #171 (harness TOLERANT_APPLY in eval/harness/common.py).
Kept in v3: Multi-hop import closure (depth 2-3): raise the universe ceiling #130 (universe ceiling, after v3 C5: discovery-union + universe/selection coverage decomposition #250 measures it), memory_pipeline re-spells the heavy phase: the #149 anti-drift fix covered only the selection half #232 (parity, with v3 prerequisites: fixes that distort the v3 measurements (§6) #245 item 5), The release-cycle equivalence gate is unreproducible off one laptop; 1150 lines of eval CI have never run #233 (eval CI, blocker for v3 data & protocol: dcbench calib/holdout split, pre-registration P1, tag paper-v3 #252 ), --budget is documented as a hard cap three times and is not one: cores bypass it (3.15x at --budget 1000) #241 / Per-file budget share: one generated data file takes 88% of a 15-file range #238 (budget prerequisites).
Moved to post-v3: Co-change mining pipeline: training/eval corpus (>=5k instances) #128 , Graph-guided BM25 query expansion (PRF) ablation #132 , Local embeddings offline ablation (hard ship gate) #133 , diffctx serve: warm daemon + incremental index (<300ms warm) #134 , Self-diff impact-oracle mode (WIP patch -> callers/tests/blast radius) #135 , Agent harness experiment: n=300, 3 arms, pre-registered kill-criterion #137 , Standalone localization benchmark on patch-derived-locations protocol #138 , Hub-suppression on/off ablation (close the v2 open question) #141 , PolyBench CST-path -> line-range resolver for fragment metrics #142 , BM25 k1 sweep + I(f) perturbation sensitivity #143 , Line-level precision/noise metrics + precision-preset decision #144 , Review-quality downstream experiment: fault-localization MRR + LLM-judge at matched budgets #166 , Residual sensitivity cells: rho/alpha blend, edge-cap K, commit signal, per-language nontrivial #169 , Churn/recency priors in relevance scoring (IR-bug-localization transfer; piggyback on cochange git-log walk) #177 , Build-graph gap-closure: go.mod, tsconfig references, JS workspaces/Nx/Turborepo + optional toolchain-exact mode + rdeps discovery #179 , Per-test coverage edges (optional instrumented run) — production/serve feature behind #134 #180 .
Product track (label v3 removed): MCP hardening: read-only, path jail, ref injection, secrets, limits, supply chain #147 , MCP publishing: official registry + directories + Claude Code plugin #148 , GEO: canonical docs, comparison page, listings, name migration, bake-off benchmark #151 , diffctx install (multi-platform MCP autoconfig) + context_savings telemetry #152 , Evaluate the auto-budget default against fixed budgets #167 .
Working thesis: v2 showed graph and lexical signals are complementary but naive mixing degrades and rank fusion loses; v3 explains why (scale incommensurability), cures it by calibrating signals into one probabilistic scale, generalises EGO/PPR as two points of one diffusion family (HKPR), and replaces the heuristic diversity bonus with DPP selection.
Order of work
paper-v3(The release-cycle equivalence gate is unreproducible off one laptop; 1150 lines of eval CI have never run #233 is a blocker here)E/Q discipline
All C1–C5 are Q-class → one eval cycle: prerequisites first, freeze, then the matrix. A Q-class change landing after the freeze invalidates calibration and forces a full rerun.
Tracker changes made 2026-08-30
TOLERANT_APPLYineval/harness/common.py).post-v3: Co-change mining pipeline: training/eval corpus (>=5k instances) #128, Graph-guided BM25 query expansion (PRF) ablation #132, Local embeddings offline ablation (hard ship gate) #133, diffctx serve: warm daemon + incremental index (<300ms warm) #134, Self-diff impact-oracle mode (WIP patch -> callers/tests/blast radius) #135, Agent harness experiment: n=300, 3 arms, pre-registered kill-criterion #137, Standalone localization benchmark on patch-derived-locations protocol #138, Hub-suppression on/off ablation (close the v2 open question) #141, PolyBench CST-path -> line-range resolver for fragment metrics #142, BM25 k1 sweep + I(f) perturbation sensitivity #143, Line-level precision/noise metrics + precision-preset decision #144, Review-quality downstream experiment: fault-localization MRR + LLM-judge at matched budgets #166, Residual sensitivity cells: rho/alpha blend, edge-cap K, commit signal, per-language nontrivial #169, Churn/recency priors in relevance scoring (IR-bug-localization transfer; piggyback on cochange git-log walk) #177, Build-graph gap-closure: go.mod, tsconfig references, JS workspaces/Nx/Turborepo + optional toolchain-exact mode + rdeps discovery #179, Per-test coverage edges (optional instrumented run) — production/serve feature behind #134 #180.v3removed): MCP hardening: read-only, path jail, ref injection, secrets, limits, supply chain #147, MCP publishing: official registry + directories + Claude Code plugin #148, GEO: canonical docs, comparison page, listings, name migration, bake-off benchmark #151, diffctx install (multi-platform MCP autoconfig) + context_savings telemetry #152, Evaluate the auto-budget default against fixed budgets #167.