B=8000. Tune on dcbench-calib; report on dcbench-holdout + public 1500. One full run, after pre-registration.
| # |
cell |
why |
| 1 |
EGO (v5) |
anchor |
| 2 |
PPR (v5) |
anchor |
| 3 |
BM25-int (v5, with gate — prerequisite 2) |
anchor; the gate changes exactly this cell |
| 4–8 |
HK t ∈ {1,3,5,10,15} |
C2 grid |
| 9 |
union (v2-style) |
fusion baseline |
| 10 |
sequential BM25→EGO |
C4 |
| 11 |
fused-sum (calibrated) |
C1 |
| 12 |
fused-noisyor |
C1 |
| 13 |
fused + DPP |
C3 |
| 14 |
discovery-union + fused |
C5 |
| 15 |
EGO @ v4 config on dcbench |
bridge to v2 |
| 16 |
fused, seeds = change_magnitude |
restored v1 signal |
- Budget curve {8k, 16k, 32k, −1}: winner and EGO only.
- Appendix ablations on calib: DPP λ-mix; calibration isotonic vs Platt vs none.
- Per-cell artifacts as in v2 (per-instance CSVs, seeds, CIs); per-role table from dcbench annotations.
- Double-run determinism check and
python -m eval equivalence on the stratified sample before the run.
B=8000. Tune on dcbench-calib; report on dcbench-holdout + public 1500. One full run, after pre-registration.
python -m eval equivalenceon the stratified sample before the run.