Repository navigation
grounded search over the docs: Hugging Face embeddings into a Joltrin vector store, with citations and I don't know - #515
Conversation
… vector store, with citations and I don't know
Gemini PR ReviewReviewed commit:
This result blocks merge. Push a fix and comment |
Gemini PR ReviewReviewed commit:
This result blocks merge. Push a fix and comment |
|
/gemini review |
Gemini PR ReviewReviewed commit:
This result blocks merge. Push a fix and comment |
|
/gemini review |
Gemini PR ReviewReviewed commit:
This result blocks merge. Push a fix and comment |
Gemini PR ReviewReviewed commit:
This result blocks merge. Push a fix and comment |
|
/gemini review |
Gemini PR ReviewReviewed commit: |
|
/gemini review |
Gemini PR ReviewReviewed commit:
|
examples/hf_grounded_search chunks five docs by heading, embeds each chunk with sentence-transformers/all-MiniLM-L6-v2, stores the vectors in a Joltrin vector store through the Python bindings (one transaction), and answers a question with the nearest passages and their file#heading citations. If the best match scores under a threshold it says it does not know.
The embedding is written with transformers and torch directly: tokenize, run the model, average the token vectors while ignoring padding, normalize to length 1. The model is pinned to one Hub commit so the results cannot move.
Retrieval only. Nothing here writes new text, so it cannot invent a sentence. Handing the passages to a language model is a separate step this does not take.
test_search.py runs real embeddings against a real store and asserts: the store's score equals the cosine similarity (0.4572 both ways on the check), all 8 questions the docs answer find their source file in the top 3, and all 6 questions about other topics are refused. The highest out-of-topic score was 0.17 and the lowest in-topic best score was 0.30, so the 0.25 threshold sits in a gap of 0.13. That is 14 questions, a starting point and not a benchmark. It also prints three near-topic probes, and two of them (a PostgreSQL replication question and a Kubernetes liveness probe question) were answered, because a score threshold cannot tell close to the topic from answered by the docs. I left that visible on purpose.
Run on an Apple silicon Mac with Python 3.12 and torch 2.14.1. A new hf-search-example workflow builds the Linux native library, installs CPU torch and runs the same test when the example, the Python bindings or the indexed docs change. It is not a required check, and I have not seen it run in CI yet.
Thanks, Gerard Recinto