The key asset
dcbench: 372 first-party instances with tier (essential/helpful) and role (definition/caller/test/config/doc/cochange) annotations, byte-exact patches, pinned repos. It resolves the main methodological weakness of a third paper on the same 1500 public instances.
Splits (frozen by commit before the first run, as in v2)
- dcbench-calib (~250, stratified by repo): C1 calibration fit, C2 t-grid, all ablations.
- dcbench-holdout (~120, never used for anything): the single confirmatory test.
- public 1500: reported for comparability with v2, with an explicit test-set-reuse disclosure paragraph (replaces "exactly once"). Calibration never touched them.
- Bonus: role annotations → per-role recall table (who finds callers, tests, config) — no related work has one.
System configuration
Only the shipped v5 (τ=0.05 + admission gate). Git tag paper-v3 before the first measurement. One bridge cell: EGO under the v4 config on dcbench, to connect with v2's numbers. The v2-paper-vs-code divergence (τ 0.12→0.05, gate) gets its own paragraph in §Reproducibility.
Pre-registration (one confirmatory test; everything else exploratory)
- P1: nontrivial_file_recall(calibrated fusion) − nontrivial(best single mode) on dcbench-holdout, B=8000. Paired one-sided permutation sign-flip; cluster bootstrap by repository (few repos → cluster CI mandatory); α=0.05; no Holm — one test.
- HKPR grid, DPP, budget curves: exploratory, CIs, no p-values in headings.
Metrics, in order (also the abstract's order)
- nontrivial_file_recall (primary) 2. file precision 3. headline recall 4. used tokens + wall-clock p50/p90 5. universe coverage (C5).
Tasks
The key asset
dcbench: 372 first-party instances with tier (essential/helpful) and role (definition/caller/test/config/doc/cochange) annotations, byte-exact patches, pinned repos. It resolves the main methodological weakness of a third paper on the same 1500 public instances.
Splits (frozen by commit before the first run, as in v2)
System configuration
Only the shipped v5 (τ=0.05 + admission gate). Git tag
paper-v3before the first measurement. One bridge cell: EGO under the v4 config on dcbench, to connect with v2's numbers. The v2-paper-vs-code divergence (τ 0.12→0.05, gate) gets its own paragraph in §Reproducibility.Pre-registration (one confirmatory test; everything else exploratory)
Metrics, in order (also the abstract's order)
Tasks
results/.../manifests)paper-v3