Benchmark-leading accuracy Β· Reduce LLM hallucinations Β· FalkorDB-fast Β· Multi-tenant Β· Graph traversal Β· 5-minute setup
Most GraphRAG systems work in demos and break under production constraints. GraphRAG SDK was built from real deployments around a simple idea: the retrieval harness matters more than the model. The result is a modular, benchmark-leading framework with predictable cost and sensible defaults that gets you from raw documents to cited answers in under 5 minutes.
It is designed to reduce LLM hallucinations: answers can be grounded in context retrieved from the knowledge graph, while applications inspect that context, validate generated claims, and gate generation when evidence is insufficient. MENTIONED_IN provenance edges trace entity mentions to their source chunks β see the Reducing LLM Hallucinations guide and the grounded answers with abstention example.
Hallucinations in RAG are usually a retrieval failure, not a model failure: the model is asked to answer from context that never contained the answer. A knowledge graph attacks that at the retrieval layer, and keeps the evidence attached to the answer.
- Relationship traversal retrieves the connected evidence vector similarity misses β multi-hop facts rarely live in one chunk that happens to be similar to the question.
- Retrieved context is traceable to source chunks;
MENTIONED_INedges trace entity mentions, andreturn_context=Truereturns the retrieval trail for application-level validation. - The ontology constrains what can be extracted, so the graph stores typed, checkable facts instead of free-form model output.
- Your application can abstain when evidence is insufficient β gate generation on retrieval and return an explicit "evidence-insufficient" response (example).
- Benchmark accuracy is the measurable outcome of all of the above β see the table below.
β Full guide: Reducing LLM hallucinations Β· API reference: Reliability and Grounding
| Rank | System | Novel (Multi-Doc) | Medical (Single-Doc) | Overall |
|---|---|---|---|---|
| 1 | FalkorDB GraphRAG SDK β | 66.09 | 76.87 | 71.48 |
| 2 | G-reasoner | 58.94 | 73.30 | 66.12 |
| 3 | AutoPrunedRetriever | 63.72 | 67.00 | 65.36 |
| 4 | HippoRAG2 | 56.48 | 64.85 | 60.67 |
| 5 | Fast-GraphRAG | 52.02 | 64.12 | 58.07 |
| 6 | RAG (w rerank) (Vector RAG) | 48.35 | 62.43 | 55.39 |
| 7 | LightRAG | 45.09 | 62.59 | 53.84 |
| 8 | HippoRAG | 44.75 | 59.08 | 51.92 |
| 9 | MS-GraphRAG (local) | 50.93 | 45.16 | 48.05 |
How these are computed. Per dataset, ACC is the unweighted mean of the four task-category scores, matching the GraphRAG-Bench leaderboard convention:
Dataset ACC = (Fact Retrieval + Complex Reasoning + Contextual Summarize + Creative Generation) / 4Overall = (Novel ACC + Medical ACC) / 2Overall is our own summary across the two datasets; the leaderboard ranks each dataset separately. Novel has 20 documents and 2,010 questions, Medical 1 corpus and 2,062 questions. FalkorDB scored August 2026 with
gpt-4o-mini(Azure OpenAI) at temperature 0.7 for both graph construction and generation,text-embedding-3-largeat 1024 dimensions, text-to-Cypher retrieval enabled, and the benchmark's owngeneration_eval.pyunmodified as the judge. Competitor numbers are from the published leaderboard, unchanged. See the GraphRAG accuracy benchmark page for the full FalkorDB vs vector RAG comparison, configuration, reproduction instructions and limitations, and the benchmark methodology page for per-category results and the full 15-system comparison.
Vectors match similar chunks. The graph traverses relationships. Every answer cites its source.
pip install graphrag-sdk[litellm]
docker run -d -p 6379:6379 -p 3000:3000 --name falkordb falkordb/falkordb:latest
export OPENAI_API_KEY="sk-..."For PDF ingestion, install the
pip install graphrag-sdk[litellm,pdf]. Ingestion sanitizes unsupported control characters in IDs and string properties before graph upserts, which helps avoid FalkorDB Cypher parse errors on noisy PDFs.
import asyncio
from graphrag_sdk import GraphRAG, ConnectionConfig, LiteLLM, LiteLLMEmbedder
async def main():
async with GraphRAG(
connection=ConnectionConfig(host="localhost", graph_name="my_graph"), # graph_name = per-tenant isolation
llm=LiteLLM(model="openai/gpt-5.5"),
embedder=LiteLLMEmbedder(model="openai/text-embedding-3-large", dimensions=256),
) as rag:
# Ingest raw text (pass a file path with the `pdf` extra installed for PDFs)
result = await rag.ingest(
text="Alice Johnson is a software engineer at Acme Corp in London.",
document_id="my_doc",
)
print(f"Nodes: {result.nodes_created}, Edges: {result.relationships_created}")
# Finalize: deduplicate entities, backfill embeddings, create indexes
await rag.finalize()
# Full RAG: retrieve + generate
answer = await rag.completion("Where does Alice work?")
print(answer.answer)
asyncio.run(main())from graphrag_sdk import GraphSchema, EntityType, RelationType
schema = GraphSchema(
entities=[
EntityType(label="Person", description="A human being"),
EntityType(label="Organization", description="A company or institution"),
EntityType(label="Location", description="A geographic location"),
],
relations=[
RelationType(label="WORKS_AT", description="Is employed by", patterns=[("Person", "Organization")]),
RelationType(label="LOCATED_IN", description="Is situated in", patterns=[("Organization", "Location")]),
],
)
async with GraphRAG(
connection=ConnectionConfig(host="localhost", graph_name="my_graph"),
llm=LiteLLM(model="openai/gpt-5.5"),
embedder=LiteLLMEmbedder(model="openai/text-embedding-3-large", dimensions=256),
schema=schema,
) as rag:
... # ingest / completion as above
β Full walkthrough: Getting Started
β Compose your own pipeline: Custom Strategies
Re-sync individual documents without rebuilding the graph. The canonical CI use case is updating the graph on PR merge β added, modified, and deleted files in one batch:
async with GraphRAG(connection=ConnectionConfig(...), llm=..., embedder=...) as graph:
result = await graph.apply_changes(
added=["docs/new_feature.md"],
modified=["docs/api.md"],
deleted=["docs/removed_page.md"],
)
await graph.finalize() # once per batch β finalize is O(graph size)
# Per-file outcomes are wrapped in BatchEntry β the batch never raises.
for entry in result.added + result.modified + result.deleted:
if not entry.is_success:
print(f"failed: {entry.error_type}: {entry.error}")The three primitives behind the wrapper:
| Method | When to use |
|---|---|
update(source, document_id=...) |
Document content changed. SHA-256 hash short-circuits no-op updates (touch-only PRs cost ~1 Cypher query). Pass if_missing="ingest" for upsert semantics. |
delete_document(document_id) |
Document removed. Cleans up entities orphaned by the deletion; preserves entities still referenced by other documents. |
apply_changes(added=..., modified=..., deleted=...) |
Heterogeneous batch. Per-file errors are collected, not raised. Does not call finalize() β caller drives that cadence. |
In file mode, document_id defaults to os.path.normpath(source) so
update("docs/x.md") matches the original ingest("docs/x.md") with
no extra plumbing. See examples/07_incremental_updates.py.
Cost model. finalize() runs cross-document deduplication, which
scans the full entity table β its cost is O(graph size), not
O(change size). Embedding backfill within finalize() is O(change
size) (only nodes/edges missing embeddings get touched). For CI use
cases, batch all PR changes through apply_changes and call
finalize once at the end of the run, not per file β per-file
finalize multiplies the dedup constant by the number of files
touched.
Crash safety. update() uses an idempotent rollforward cutover:
the new content is written to a __pending__ Document, then a single
atomic Cypher statement marks ready_to_commit=true, then the live
document is replaced. A crash before the marker discards the pending
on retry; a crash after the marker rolls forward to completion. Either
way, retrying the same update() call is safe and converges on the
correct final state.
Concurrency. apply_changes exposes two knobs: max_concurrency
(adds, default 3) and update_concurrency (modifies, default 1).
Updates default to 1 because orphan-cleanup correctness under
concurrent updates depends on a pipeline-ordering invariant; raising
that default is safe only if you've verified your concurrent updates
can never share an entity. The integration test
test_concurrent_updates_preserve_shared_entity is the tripwire that
guards the default.
| Area | Step | Cost |
|---|---|---|
| Ingestion | Extract entities & relations | LLM |
| Ingestion | Resolve & deduplicate entities | LLM |
| Ingestion | Embed & index | LLM |
| Retrieval | Vector search | DB |
| Retrieval | Full-text search | DB |
| Retrieval | Text-to-Cypher (experimental) | LLM |
| Retrieval | Cypher queries | DB |
| Retrieval | Relationship expansion | DB |
| Retrieval | Cosine reranking | Local |
π‘ Retrieved context can be traced to source chunks;
MENTIONED_INedges connect entity mentions to chunks. Passreturn_context=Truetocompletion()so your application can inspect the retrieval trail and validate generated claims.
Working starters β clone, plug in your source, ship.
| # | Example | What you'll build |
|---|---|---|
| 1 | Quick Start | Your first ingest-and-query loop in under 30 lines |
| 2 | PDF with Schema | A PDF Q&A bot with your own entity and relation types |
| 3 | Custom Strategies | Composing ingestion strategies explicitly |
| 4 | Custom Provider | Plug in any LLM or embedder behind a clean interface |
| 5 | Notebook Demo | An interactive walkthrough that shows the provenance trail |
| 6 | Markdown, Document-Aware | Structure-preserving Markdown ingestion with queryable heading breadcrumbs |
| 7 | Incremental Updates | update, delete_document, and apply_changes for CI-driven graph syncs |
| 8 | Ontology Lifecycle | Declare an ontology, ingest with it, and round-trip it as JSON config |
| 9 | Ontology Evolution | Mutating schema evolution β rename types and atomically add attributes with LLM backfill |
| 10 | Ontology Discovery | Discover an ontology from raw sources and propose extensions as new docs arrive |
| β | Grounded Answers with Abstention | Cite the retrieved context behind an answer, and abstain when the graph has no supporting evidence |
Full documentation: https://docs.falkordb.com/graphrag
| Guide | Description |
|---|---|
| Getting Started | Step-by-step tutorial from install to first query |
| Architecture | Pipeline design, graph schema, retrieval strategy |
| Reducing LLM Hallucinations | Grounded retrieval, source provenance, and abstention |
| Configuration | Connection, providers, and tuning reference |
| Strategies | All ABCs and built-in implementations |
| Providers | LLM and embedder configuration guide |
| Reliability and Grounding | Grounding, provenance and abstention mapped to the APIs that implement them |
| Benchmark | Methodology, results, and reproduction instructions |
| Accuracy Benchmark: FalkorDB vs Vector RAG | 71.48 vs 55.39 comparison, evaluation definition, limitations, how to cite |
| API Reference | Full API documentation |
- 2024-06: First public release
- 2024-Q4: PDF ingestion and multi-provider LLMs
- 2025-Q1βQ2: Pluggable providers and pipeline tuning
- 2025-Q3: Sharper retrieval, deeper test coverage
- π 2026-04: Version 1.0 is released with a new set of benchmarks based on a year's worth of research and customer PoCs
- π¦ Still on the v0.x API? Pin the legacy release:
pip install graphrag-sdk==0.8.2
- π¦ Still on the v0.x API? Pin the legacy release:
- 2026-Q2: Production observability; expand ingestion support β tables, structured data
- 2026-Q3: Introduce Agentic GraphRAG; complete PDF ingestion
- 2026-Q4: Smarter retrieval β dynamic traversal, temporal graph
We welcome contributions! See CONTRIBUTING.md for development setup, testing, and code style guidelines.
Please read our Code of Conduct before participating.
- Discord -- Ask questions, share what you build
- GitHub Discussions -- Feature ideas, Q&A
- Issues -- Bug reports and feature requests
If you use GraphRAG SDK in your research, please cite:
@software{graphrag_sdk,
title = {GraphRAG SDK: A Modular Graph RAG Framework},
author = {FalkorDB},
year = {2026},
url = {https://github.com/FalkorDB/GraphRAG-SDK},
}
