Skip to content

feat(lancedb): build an IVF_FLAT index on vector columns - #461

Merged
gloryfromca merged 7 commits into
mainfrom
feat/vector-index
Sep 24, 2026
Merged

gloryfromca merged 7 commits into
mainfrom
feat/vector-index

Conversation

@gloryfromca

@gloryfromca gloryfromca commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Summary

EverOS built FTS indexes on its LanceDB tables and nothing on the vector columns, so every nearest_to was a brute-force scan of the whole column — linear in rows and in bytes. Measured on the Windows soak box (1024-dim float32): 10k rows = 42 MB = 155 ms per scan, 27k rows = 112 MB = 590 ms; a hybrid search issues two or three of them, which is where its 0.5–1.6 s idle latency went and why it climbed through the 10-hour soak (PR #454).

This adds BaseLanceTable.ensure_vector_indexes(table, *, min_rows): one IVF_FLAT, cosine index (matching every nearest_to's distance_type("cosine")) per vector column once it holds [lancedb] vector_index_min_rows (default 2000) non-null vectors. All-null columns (a Tier 1 store has no embeddings) and small stores are left alone. Two cases do work, everything else is a no-op:

  • no index yet and the column crossed the threshold → build one;
  • the index has grown delta indices → retrain it in place. Every optimize() on a table with new rows appends a delta IVF index and never merges it (num_indices +1 per light beat, confirmed locally and on the soak box: 35 deltas after 25 minutes). A query probes every delta, so latency climbed with the beats since the last rebuild — 27k × 1024 rows: 5.9 ms at 0 deltas, 25.7 ms at 100, 303 ms at 400, worse than the 24 ms scan the index replaces. create_index(replace=True) collapses them (optimize(retrain=True) does not).

It runs at startup next to the FTS pass, on the cascade's heavy beat (every 300 s, so at most ~30 deltas accumulate under sustained writes), and on the 12 h rebuild cadence (which used to drop every non-BM25 index and would have deleted the vector indexes; it now keeps them and only retrains when deltas exist).

The heavy-beat call reaches the LanceDB layer through the IndexRepository protocol: forwarded by RoutedIndexRepository and LanceIndexRepository, a documented no-op on Milvus, called unguarded by the worker. (The first version used getattr(repo, "ensure_vector_indexes", None) on the routed repository, which did not forward it — dead code, caught in review.)

Query side: VECTOR_INDEX_ROWS_PER_PARTITION = 4096 and VECTOR_QUERY_NPROBES = 32 live together in core/persistence/lancedb/base.py; every nearest_to sets .nprobes(...), so the search is exact up to ~130k rows and degrades gradually past it instead of depending on lance's changing defaults.

Known ceiling (ponytail: in the docstring): the collapse is a full retrain per heavy beat under sustained writes; merging deltas needs pylance's optimize_indices, which is not a dependency. Revisit past ~500k rows.

Before / after

Lance-level scan vs indexed query on the 27k-row run1 tables and the review findings: see the comments. Server-level: the Windows soak on the integration branch before this revision (unfixed deltas) is the "before"; the same soak on the revised code is running and its numbers follow in a comment.

Area

  • architecture (storage)

Verification

  • tests/unit/test_core/test_persistence/test_lancedb/test_vector_index.py (real LanceDB in tmp_path): vector columns detected from the Arrow schema; nothing below the threshold; nothing on all-null vectors; IvfFlat built at the threshold; delta indexes left by two add + optimize() beats are collapsed to one (the precondition num_indices == 3 is asserted, so a LanceDB that starts merging on its own fails the precondition rather than passing silently); nprobes accepted on an unindexed column; indexed top-5 == brute-force top-5.
  • tests/unit/test_infra/test_index_router_maintenance.py: every repository class exposes the four maintenance methods, the protocol lists them, the router forwards ensure_vector_indexes.
  • tests/unit/test_memory/test_cascade/test_worker.py::test_heavy_beat_ensures_vector_indexes: heavy beat calls it, light beat does not.
  • Each new test fails under its own mutation (no delta retrain → [] == ['vector']; router forwarding removed → method missing; worker call removed → 0 calls). test_core/test_persistence + test_infra + test_cascade: 1004 passed. ruff + lint-imports green.

Checklist

  • tests added, mutation checked
  • setting documented in default.toml
  • adversarial review round 1 addressed point by point (comment); round 2 in progress

Notes for Reviewers

The threshold is a setting because the break-even depends on the disk: on the laptop SSD the scan already cost 155 ms at 10k rows; a NVMe workstation can afford more before indexing. The L2 LanceRepoBase.search path (tests only) returns cosine neighbours once an index exists; left as is and noted in the review reply.

🤖 Generated with Claude Code

@gloryfromca

Copy link
Copy Markdown
Member Author

Before/after on the soak box, Lance-level (a copy of the run1 tables, 20 real row vectors as queries, best of 3, select(["id"]).limit(10), cosine). The box was running another job at the time, so absolute numbers are pessimistic; the ratios and recall are what matters.

table rows brute-force scan p50 IVF_FLAT p50 speed-up recall@10 vs scan top-1 index build
episode 27 238 245 ms 84 ms ×2.9 0.980 1.00 2.1 s
atomic_fact 26 715 244 ms 94 ms ×2.6 0.965 1.00 1.8 s

For reference, IVF_PQ on the same data: 9–11 ms (×22–27) but recall@10 collapses to 0.18–0.21 without a query-side refine_factor (top-1 still 1.00). This PR keeps IVF_FLAT: same neighbours as the scan, no query change. If the remaining ~90 ms matters, the next step is IVF_PQ plus refine_factor in dense_search (re-rank the candidates with exact vectors), measured on real embeddings rather than the soak's synthetic ones.

Server-level hybrid /search before/after will follow once the running soak has released the box.

EverOS built FTS indexes on its LanceDB tables and nothing on the vector
columns, so every `nearest_to` was a brute-force scan of the whole column:
linear in rows and bytes. Measured on the Windows soak box (1024-dim
float32): 10k rows = 42 MB = 155 ms per scan, 27k rows = 112 MB = 590 ms;
a hybrid search issues two or three scans, which is where its 0.5–1.6 s
idle latency went and why it climbed through the 10-hour soak.

`BaseLanceTable.ensure_vector_indexes` creates an IVF_FLAT (cosine, to
match `dense_search`) index on each vector column once it holds
`[lancedb] vector_index_min_rows` (default 2000) non-null vectors; all-null
columns (a Tier 1 store) and small stores are left alone. It runs at
startup with the FTS pass, on the cascade's heavy maintenance beat (so a
table that crosses the threshold while the server runs gets indexed), and
with `replace=True` on the rebuild cadence. `rebuild_indexes` used to drop
every index not on a BM25 column — it now keeps the vector indexes it would
otherwise have removed every 12 h. LanceDB's `optimize()` merges new rows
into an existing index, so the unindexed tail stays small between beats.

Tests: columns detected from the Arrow schema; nothing below the threshold
or on all-null vectors; IvfFlat built at the threshold; idempotent unless
replaced; indexed top-5 equals the brute-force top-5.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@gloryfromca gloryfromca changed the title feat(lancedb): build an IVF_FLAT index on vector columns past a row threshold feat(lancedb): build an IVF_FLAT index on vector columns Sep 24, 2026
Review of the first version (adversarial subagent, probes on lancedb 0.34)
found two blockers:

* Every optimize() on a table with new rows appends a delta to the IVF
  index and never merges it (num_indices +1 per light beat, confirmed
  locally). A query probes every delta, so latency climbed with the beats
  since the last rebuild: 27k x 1024 rows measured 5.9 ms at 0 deltas,
  25.7 ms at 100, 303 ms at 400 -- worse than the 24 ms scan the index
  replaces. Only create_index(replace=True) collapses them
  (optimize(retrain=True) does not); ensure_vector_indexes now retrains a
  column whose index has num_indices > 1, and the heavy beat (300 s) calls
  it, so at most ~30 deltas accumulate under sustained writes.
* The heavy-beat hook was dead code: the worker holds the routed
  repository, which did not forward ensure_vector_indexes, and the
  getattr guard skipped silently. The method is now on the IndexRepository
  protocol, forwarded by the router and the LanceDB backend, a no-op on
  Milvus, and the worker calls it unguarded.

Also: pin the IVF partition size (4096 rows) and nprobes (32) as one pair
of constants used by every nearest_to, so the search is exact up to ~130k
rows instead of depending on lance's changing defaults; drop the replace=
parameter (the conditional covers the rebuild cadence and removes the
retrain-at-startup the review also flagged); fix the read-timeout
docstring that still said no ANN index is built.

Tests: delta collapse (real LanceDB; precondition asserts one delta per
beat), nprobes accepted on an unindexed column, every repository class
exposes the maintenance methods and the router forwards them, the heavy
beat calls ensure_vector_indexes. Each fails under its own mutation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@gloryfromca

Copy link
Copy Markdown
Member Author

Revision pushed after an adversarial review pass (read-only subagent with LanceDB probes). Findings and what changed:

# finding change
1 blocker every light-beat optimize() leaves one more IVF delta index, never merged; query time grows with delta count (27k×1024: 5.9 ms @0 → 25.7 ms @100 → 303 ms @400, vs 24 ms flat scan) ensure_vector_indexes retrains a column whose index has num_indices > 1 (create_index(replace=True); optimize(retrain=True) does not collapse). Heavy beat calls it, so ≤ ~30 deltas accumulate. Test asserts the precondition (one delta per beat) and the collapse.
2 blocker heavy-beat hook was dead code: the worker holds the routed repo, which did not forward the method; getattr guard skipped silently ensure_vector_indexes added to the IndexRepository protocol, forwarded by router + LanceDB backend, no-op on Milvus, called unguarded. Tests: every repository class exposes the maintenance methods; router forwards; heavy beat calls it.
3 should-fix recall depends on lance's default partition count; nprobes unset VECTOR_INDEX_ROWS_PER_PARTITION = 4096, VECTOR_QUERY_NPROBES = 32 in one place, used by every nearest_to; exact up to ~130k rows, graceful past it.
4 should-fix startup build then the rebuild loop retrains immediately gone with the conditional: rebuild_indexes only retrains when deltas exist.
6 nit read-timeout docstring said no ANN index is built fixed.
5 nit the L2 LanceRepoBase.search path returns cosine neighbours once an index exists left as is: only tests call it; noted here.

Known ceiling (marked ponytail: in the docstring): the collapse is a full retrain per heavy beat under sustained writes; merging deltas needs pylance's optimize_indices, not a dependency. Revisit past ~500k rows.

gloryfromca and others added 2 commits September 24, 2026 16:11
Second review pass: with the cap at one delta, a trickle writer (one small
upsert per 5-minute window) paid a full index rewrite every heavy beat --
107 MB per column at 27k x 1024 rows, 391 MB at 100k -- while 16 deltas
cost 1-3 ms extra per query. Under sustained writes the light beat adds
~30 deltas per heavy beat, so the cap changes nothing there; it only spares
light users the write amplification.

Also cover the gap the first review hit one layer down: the LanceDB backend
forwarding is asserted with a stub (a 'pass' body satisfied the protocol
check alone), and the docstring names the cross-process optimize() commit
conflict that can preempt a retrain (same exposure as prune; retried on
the next heavy beat).

Tests: the cap test is monkeypatched to 2 and asserts both sides (2 deltas
left alone, 3 collapsed); it fails when the cap is ignored. The forwarding
test fails with a 'pass' body.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@gloryfromca

Copy link
Copy Markdown
Member Author

Second adversarial pass on the revision: no blocker. Confirmed by probes on lancedb 0.34: the collapse works at the soak's shape (27k×1024, 30 beats: num_indices 31→1, query 7.9→5.8 ms; retrain 0.5 s / 107 MB, 1.4 s / 391 MB at 100k), create_index(replace=True) is one atomic swap (index absent 0/450 polls during 3 retrains under 8 query loops + a writer, 0 errors, 0 empties), the pinned partitions/nprobes are what make recall exact at 100k (1.000 vs 0.885 with lance's defaults), and the protocol change breaks no backend or test fake.

Two should-fixes, both in the follow-up commit:

finding change
retrain at num_indices > 1 = a full index rewrite per heavy beat even for a trickle writer (up to ~225 GB/day at 100k rows × 2 columns under sustained writes) VECTOR_INDEX_MAX_DELTAS = 16: 16 deltas cost 1–3 ms per query; under sustained writes ~30 deltas arrive per heavy beat so the steady state is unchanged, light users retrain 16× less. Test asserts both sides of the cap.
the maintenance-method test checked attribute presence only; a pass body on LanceIndexRepository.ensure_vector_indexes would have passed forwarding asserted with a stub repo; fails on a pass body.

Noted, not changed: a concurrent optimize() from another process (cascade sync / backfill) can preempt the retrain with the benign CreateIndex preempted by CreateIndex conflict; the next heavy beat retries (same exposure prune has) — now in the docstring. Lock hold at ~500k rows (≈14 s retrain vs the 15 s write deadline) is the ceiling the ponytail: note points at.

ensure_business_indexes runs in every process that opens the root: the
server lifespan and the CLI's _runtime. With the vector step in it, a
read-only 'everos cascade status' trained an IVF index on the running
server's table the moment it crossed the row threshold and failed with
'Retryable commit conflict' (Windows soak, final run: 10/9 storm errors in
15 minutes; the first two runs of this PR had none because the CLI storms
ran before any table reached 2000 rows).

Vector indexes belong to the cascade worker alone: its first rebuild sweep
at server start builds a missing one (rebuild_indexes -> ensure_vector_
indexes) and the heavy beat keeps it healthy. FTS stays in the startup
pass because a search on a column without its inverted index raises.

Test: a spy on BaseLanceTable.ensure_vector_indexes must see no call from
ensure_business_indexes; re-adding the call fails it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@gloryfromca

Copy link
Copy Markdown
Member Author

Server-level check of the delta collapse, same soak harness, same box, same ~3.5k-row point (ceiling load, 20 min in):

code episode vector_idx num_indices atomic_fact foresight
first version (deltas never collapsed) 35 23 19
revision (heavy beat retrains past the cap) 3 4 –

The hybrid /search p50 curve over the first 25 minutes is the same shape in both runs (≈440 → 600 → 780 → 880 → 1050 ms) and the keyword-only curve climbs with it, so at this table size the latency growth is the server-side load/table growth seen in run 2, not the delta count — the index keeps its win (Lance-level 245 → 84 ms) without the delta tail eating it later.

Also found by this soak and fixed in 669b23d: the startup hook lived in ensure_business_indexes, which the CLI's _runtime also calls, so a read-only cascade status trained the index on the live server's table and hit Retryable commit conflict (see the commit message). Vector-index maintenance is now the cascade worker's alone.

@gloryfromca gloryfromca reopened this Sep 24, 2026
@gloryfromca
gloryfromca enabled auto-merge (squash) September 24, 2026 08:54
@gloryfromca

Copy link
Copy Markdown
Member Author

Existing-store check on the Windows box (code 883007c = this PR + #462 + #463 on main), against a copy of a real 3.9k-row store from the soak, driving the product's own calls:

store state call result
pre-#461 store, no vector index episode_repo.rebuild_indexes() (the worker's first sweep at server start) both IvfFlat indexes built in 1.7 s; the flat search before it answered in 79 ms, nothing blocked
19 delta indices after 18 add+optimize beats episode_repo.ensure_vector_indexes() (the heavy beat) collapsed to 1 in 2.7 s
healthy indexed store rebuild_indexes() again 0.5 s, indexes kept, nearest_to returns the seed row

So an install upgrading onto this PR gets its index on the first server start (seconds; searches keep working flat until then), and the delta cap is enforced on real data, not only on the unit test's 60 rows.

Pins what the code already does: an exception from ensure_vector_indexes
on the heavy beat is caught with the other maintenance failures, counted
toward optimize_failures (so the health signal can show a streak), leaves
the prune that ran before it credited, and the next beat still runs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@gloryfromca
gloryfromca merged commit 4e905dc into main Sep 24, 2026
10 checks passed
@gloryfromca
gloryfromca deleted the feat/vector-index branch September 24, 2026 09:18
This was referenced Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants