Skip to content

v3 data & protocol: dcbench calib/holdout split, pre-registration P1, tag paper-v3 #252

Description

@nikolay-e

The key asset

dcbench: 372 first-party instances with tier (essential/helpful) and role (definition/caller/test/config/doc/cochange) annotations, byte-exact patches, pinned repos. It resolves the main methodological weakness of a third paper on the same 1500 public instances.

Splits (frozen by commit before the first run, as in v2)

  • dcbench-calib (~250, stratified by repo): C1 calibration fit, C2 t-grid, all ablations.
  • dcbench-holdout (~120, never used for anything): the single confirmatory test.
  • public 1500: reported for comparability with v2, with an explicit test-set-reuse disclosure paragraph (replaces "exactly once"). Calibration never touched them.
  • Bonus: role annotations → per-role recall table (who finds callers, tests, config) — no related work has one.

System configuration

Only the shipped v5 (τ=0.05 + admission gate). Git tag paper-v3 before the first measurement. One bridge cell: EGO under the v4 config on dcbench, to connect with v2's numbers. The v2-paper-vs-code divergence (τ 0.12→0.05, gate) gets its own paragraph in §Reproducibility.

Pre-registration (one confirmatory test; everything else exploratory)

  • P1: nontrivial_file_recall(calibrated fusion) − nontrivial(best single mode) on dcbench-holdout, B=8000. Paired one-sided permutation sign-flip; cluster bootstrap by repository (few repos → cluster CI mandatory); α=0.05; no Holm — one test.
  • HKPR grid, DPP, budget curves: exploratory, CIs, no p-values in headings.

Metrics, in order (also the abstract's order)

  1. nontrivial_file_recall (primary) 2. file precision 3. headline recall 4. used tokens + wall-clock p50/p90 5. universe coverage (C5).

Tasks

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions