feat(lancedb): build an IVF_FLAT index on vector columns - #461
Conversation
|
Before/after on the soak box, Lance-level (a copy of the run1 tables, 20 real row vectors as queries, best of 3,
For reference, IVF_PQ on the same data: 9–11 ms (×22–27) but recall@10 collapses to 0.18–0.21 without a query-side Server-level hybrid |
EverOS built FTS indexes on its LanceDB tables and nothing on the vector columns, so every `nearest_to` was a brute-force scan of the whole column: linear in rows and bytes. Measured on the Windows soak box (1024-dim float32): 10k rows = 42 MB = 155 ms per scan, 27k rows = 112 MB = 590 ms; a hybrid search issues two or three scans, which is where its 0.5–1.6 s idle latency went and why it climbed through the 10-hour soak. `BaseLanceTable.ensure_vector_indexes` creates an IVF_FLAT (cosine, to match `dense_search`) index on each vector column once it holds `[lancedb] vector_index_min_rows` (default 2000) non-null vectors; all-null columns (a Tier 1 store) and small stores are left alone. It runs at startup with the FTS pass, on the cascade's heavy maintenance beat (so a table that crosses the threshold while the server runs gets indexed), and with `replace=True` on the rebuild cadence. `rebuild_indexes` used to drop every index not on a BM25 column — it now keeps the vector indexes it would otherwise have removed every 12 h. LanceDB's `optimize()` merges new rows into an existing index, so the unindexed tail stays small between beats. Tests: columns detected from the Arrow schema; nothing below the threshold or on all-null vectors; IvfFlat built at the threshold; idempotent unless replaced; indexed top-5 equals the brute-force top-5. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
104ccdf to
2a0545e
Compare
Review of the first version (adversarial subagent, probes on lancedb 0.34) found two blockers: * Every optimize() on a table with new rows appends a delta to the IVF index and never merges it (num_indices +1 per light beat, confirmed locally). A query probes every delta, so latency climbed with the beats since the last rebuild: 27k x 1024 rows measured 5.9 ms at 0 deltas, 25.7 ms at 100, 303 ms at 400 -- worse than the 24 ms scan the index replaces. Only create_index(replace=True) collapses them (optimize(retrain=True) does not); ensure_vector_indexes now retrains a column whose index has num_indices > 1, and the heavy beat (300 s) calls it, so at most ~30 deltas accumulate under sustained writes. * The heavy-beat hook was dead code: the worker holds the routed repository, which did not forward ensure_vector_indexes, and the getattr guard skipped silently. The method is now on the IndexRepository protocol, forwarded by the router and the LanceDB backend, a no-op on Milvus, and the worker calls it unguarded. Also: pin the IVF partition size (4096 rows) and nprobes (32) as one pair of constants used by every nearest_to, so the search is exact up to ~130k rows instead of depending on lance's changing defaults; drop the replace= parameter (the conditional covers the rebuild cadence and removes the retrain-at-startup the review also flagged); fix the read-timeout docstring that still said no ANN index is built. Tests: delta collapse (real LanceDB; precondition asserts one delta per beat), nprobes accepted on an unindexed column, every repository class exposes the maintenance methods and the router forwards them, the heavy beat calls ensure_vector_indexes. Each fails under its own mutation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Revision pushed after an adversarial review pass (read-only subagent with LanceDB probes). Findings and what changed:
Known ceiling (marked |
Second review pass: with the cap at one delta, a trickle writer (one small upsert per 5-minute window) paid a full index rewrite every heavy beat -- 107 MB per column at 27k x 1024 rows, 391 MB at 100k -- while 16 deltas cost 1-3 ms extra per query. Under sustained writes the light beat adds ~30 deltas per heavy beat, so the cap changes nothing there; it only spares light users the write amplification. Also cover the gap the first review hit one layer down: the LanceDB backend forwarding is asserted with a stub (a 'pass' body satisfied the protocol check alone), and the docstring names the cross-process optimize() commit conflict that can preempt a retrain (same exposure as prune; retried on the next heavy beat). Tests: the cap test is monkeypatched to 2 and asserts both sides (2 deltas left alone, 3 collapsed); it fails when the cap is ignored. The forwarding test fails with a 'pass' body. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Second adversarial pass on the revision: no blocker. Confirmed by probes on lancedb 0.34: the collapse works at the soak's shape (27k×1024, 30 beats: Two should-fixes, both in the follow-up commit:
Noted, not changed: a concurrent |
ensure_business_indexes runs in every process that opens the root: the server lifespan and the CLI's _runtime. With the vector step in it, a read-only 'everos cascade status' trained an IVF index on the running server's table the moment it crossed the row threshold and failed with 'Retryable commit conflict' (Windows soak, final run: 10/9 storm errors in 15 minutes; the first two runs of this PR had none because the CLI storms ran before any table reached 2000 rows). Vector indexes belong to the cascade worker alone: its first rebuild sweep at server start builds a missing one (rebuild_indexes -> ensure_vector_ indexes) and the heavy beat keeps it healthy. FTS stays in the startup pass because a search on a column without its inverted index raises. Test: a spy on BaseLanceTable.ensure_vector_indexes must see no call from ensure_business_indexes; re-adding the call fails it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Server-level check of the delta collapse, same soak harness, same box, same ~3.5k-row point (ceiling load, 20 min in):
The hybrid Also found by this soak and fixed in 669b23d: the startup hook lived in |
|
Existing-store check on the Windows box (code 883007c = this PR + #462 + #463 on main), against a copy of a real 3.9k-row store from the soak, driving the product's own calls:
So an install upgrading onto this PR gets its index on the first server start (seconds; searches keep working flat until then), and the delta cap is enforced on real data, not only on the unit test's 60 rows. |
Pins what the code already does: an exception from ensure_vector_indexes on the heavy beat is caught with the other maintenance failures, counted toward optimize_failures (so the health signal can show a streak), leaves the prune that ran before it credited, and the next beat still runs. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Summary
EverOS built FTS indexes on its LanceDB tables and nothing on the vector columns, so every
nearest_towas a brute-force scan of the whole column — linear in rows and in bytes. Measured on the Windows soak box (1024-dim float32): 10k rows = 42 MB = 155 ms per scan, 27k rows = 112 MB = 590 ms; a hybrid search issues two or three of them, which is where its 0.5–1.6 s idle latency went and why it climbed through the 10-hour soak (PR #454).This adds
BaseLanceTable.ensure_vector_indexes(table, *, min_rows): one IVF_FLAT, cosine index (matching everynearest_to'sdistance_type("cosine")) per vector column once it holds[lancedb] vector_index_min_rows(default 2000) non-null vectors. All-null columns (a Tier 1 store has no embeddings) and small stores are left alone. Two cases do work, everything else is a no-op:optimize()on a table with new rows appends a delta IVF index and never merges it (num_indices+1 per light beat, confirmed locally and on the soak box: 35 deltas after 25 minutes). A query probes every delta, so latency climbed with the beats since the last rebuild — 27k × 1024 rows: 5.9 ms at 0 deltas, 25.7 ms at 100, 303 ms at 400, worse than the 24 ms scan the index replaces.create_index(replace=True)collapses them (optimize(retrain=True)does not).It runs at startup next to the FTS pass, on the cascade's heavy beat (every 300 s, so at most ~30 deltas accumulate under sustained writes), and on the 12 h rebuild cadence (which used to drop every non-BM25 index and would have deleted the vector indexes; it now keeps them and only retrains when deltas exist).
The heavy-beat call reaches the LanceDB layer through the
IndexRepositoryprotocol: forwarded byRoutedIndexRepositoryandLanceIndexRepository, a documented no-op on Milvus, called unguarded by the worker. (The first version usedgetattr(repo, "ensure_vector_indexes", None)on the routed repository, which did not forward it — dead code, caught in review.)Query side:
VECTOR_INDEX_ROWS_PER_PARTITION = 4096andVECTOR_QUERY_NPROBES = 32live together incore/persistence/lancedb/base.py; everynearest_tosets.nprobes(...), so the search is exact up to ~130k rows and degrades gradually past it instead of depending on lance's changing defaults.Known ceiling (
ponytail:in the docstring): the collapse is a full retrain per heavy beat under sustained writes; merging deltas needs pylance'soptimize_indices, which is not a dependency. Revisit past ~500k rows.Before / after
Lance-level scan vs indexed query on the 27k-row run1 tables and the review findings: see the comments. Server-level: the Windows soak on the integration branch before this revision (unfixed deltas) is the "before"; the same soak on the revised code is running and its numbers follow in a comment.
Area
Verification
tests/unit/test_core/test_persistence/test_lancedb/test_vector_index.py(real LanceDB intmp_path): vector columns detected from the Arrow schema; nothing below the threshold; nothing on all-null vectors;IvfFlatbuilt at the threshold; delta indexes left by twoadd+optimize()beats are collapsed to one (the preconditionnum_indices == 3is asserted, so a LanceDB that starts merging on its own fails the precondition rather than passing silently);nprobesaccepted on an unindexed column; indexed top-5 == brute-force top-5.tests/unit/test_infra/test_index_router_maintenance.py: every repository class exposes the four maintenance methods, the protocol lists them, the router forwardsensure_vector_indexes.tests/unit/test_memory/test_cascade/test_worker.py::test_heavy_beat_ensures_vector_indexes: heavy beat calls it, light beat does not.[] == ['vector']; router forwarding removed → method missing; worker call removed → 0 calls).test_core/test_persistence+test_infra+test_cascade: 1004 passed.ruff+lint-importsgreen.Checklist
default.tomlNotes for Reviewers
The threshold is a setting because the break-even depends on the disk: on the laptop SSD the scan already cost 155 ms at 10k rows; a NVMe workstation can afford more before indexing. The L2
LanceRepoBase.searchpath (tests only) returns cosine neighbours once an index exists; left as is and noted in the review reply.🤖 Generated with Claude Code