A live Java/Python code map for developers and AI agents.
Symbols, call graph, references, impact, dataflow and taint with explicit coverage and uncertainty — local-first, staleness-aware, model-agnostic.
Quick start · Why it exists · Product contract · Documentation · When to use it · Benchmarks · Maturity · Roadmap
An AI coding agent asked "what breaks if I change this function?" has two bad options: read dozens of files (slow, expensive, easy to miss a caller) or trust an index that was built at some point in the past and may already be wrong. In a live editing session — the moment the answer matters most — most code indexes are stale, and a stale graph answers confidently and incorrectly. That single failure mode is the number-one risk in this space.
GraphCodeMap parses your repository with tree-sitter into a SQLite-backed graph and exposes it as focused tools — a library, a CLI, and an MCP server any agent can call. It is built around one invariant:
The code is the source of truth; the graph is a derived cache.
Two consequences make it trustworthy where other indexes are not:
- Freshness is checked at answer time. Each row carries the exact content hash of the file it came from. Relevant paths use read-repair, while watcher and full-sweep backstops discover new/deleted files. Size and mtime are hints; unchanged metadata never substitutes for content verification.
- Every fact declares its confidence. Call edges are labeled
certain,inferred, orpossible. Static-analysis limits are stated, never hidden. Transitive queries propagate the minimum confidence along the path.
This is epistemic honesty as an engineering principle — and it is the whole
point. An agent that can trust a certain answer stops re-verifying by reading
files, which is where the graph turns into both a correctness win and a token
win.
Product reset: Java and Python are the phase-one focus. Other extractors remain experimental compatibility surfaces until they pass the same product gates. The canonical scope, graph vocabulary, acceptance criteria and current gaps are in the Product Contract.
git clone https://github.com/Victor-Alves0/graphcodemap.git
cd graphcodemap
python -m pip install -e ".[mcp,l1]"
codegraph setup --install # detect languages; show + confirm a pinned plan
codegraph index . --l1 # build the index and promote semantic edges
codegraph doctor # verify resolver health and % certain
codegraph overview # ranked map of the repo (PageRank)
codegraph tree # physical folders/files, hashes and index state
codegraph history # Git-aware graph and analysis-stage revisions
codegraph semantic-coverage # certain/fallback/unresolved callsites and why
codegraph find validate_token # locate symbols
codegraph impact auth.TokenService.validate # what breaks if I change this?
codegraph callers auth.TokenService.validate # who calls it (with confidence)
codegraph taint --entry handle_request # untrusted input → dangerous sink
codegraph dataflow-build # persist Java/Python flows_to
codegraph flow-path pkg.handle.request pkg.save.value # reusable value path
codegraph path-traversal # external sources → CWE-22 sinks
codegraph path-traversal app.download # or explicit entry parameters
codegraph visualize --mode impact --symbol validate_token # investigate as HTMLPoint any MCP-capable agent at your repo:
codegraph mcp --install # prepares MCP + repo languages, then startsThe distribution is not published on PyPI yet, so the documented path is an
editable install from this checkout. The core remains useful without L1, but
semantic call edges cannot become certain until their resolver is ready.
codegraph setup first reuses installed/repo-local tools, then offers an
explicit versioned installation; it never downloads merely because a repo was
opened. python -m pip install -e ".[l1]" is the direct Python/Jedi shortcut.
See Languages & Resolvers.
Or embed it as a library — the importable package is codegraph (like pillow→PIL):
from codegraph import CodeGraph
cg = CodeGraph(".")
cg.index()
rows, env = cg.find_symbol("validate")New here? Start with Getting Started.
| Question | Tool |
|---|---|
| Where does this repo begin? | overview — ranked map by PageRank |
| Where is this symbol? | find, info |
| Who calls this / what does it call? | callers, callees |
| What breaks if I change this? | impact, change-impact, affected-modules |
| Which tests cover this? | related-tests |
| Where does untrusted input flow? | dataflow, taint, reaches |
| Can external input reach a path/file API? | path-traversal (repo-wide or explicit entry) |
| What are the subsystems here? | communities |
| Where is every file, and is its graph current? | tree |
| Which repository/analysis revision produced this graph? | history |
| Which calls are semantically proven, and why not the rest? | semantic-coverage |
| What should I read for this task? | suggest, explain |
| Show me the neighborhood | visualize — seeded, interactive HTML |
Full reference: CLI · MCP tools · Library API · Java analysis contract.
- Read-repair freshness guarantee. A watcher keeps the index hot, a boot scan catches offline changes, and every query verifies content-hashes before answering. Measured to 100k+ files. → Concepts
- Confidence-typed edges.
certain/inferred/possible, withcertaincoming from real semantic resolution (L1). → Concepts - Impact & change analysis. Transitive reverse-reachability, ranked by PageRank × path-confidence; feed it a git diff to ask "what does my branch break?" → CLI
- Dataflow & taint (CPG-lite). Java/Python def-use and interprocedural
flows_toare persisted with stable nodes and per-function input hashes;flow-pathqueries them without reparsing. The first G4 rule,path-traversal, now consumes only those persisted paths.call_resultnodes let configured/framework sources seed a repo-wide scan and let sanitizer transformations cut paths; an explicit entry-parameter mode remains available. Absence isunknown, never a safety proof. The broader compatibility security engine remains flow-sensitive in 18 of the 19 dedicated code languages: a redefinition kills the taint, sox = input(); x = escape(x); sink(x)is correctly reported clean. → Concepts - Semantic L1 via LSP. Promotes edges to
certainthrough one generic LSP client; every dedicated language has a resolver wired. → Languages & Resolvers - Agent-oriented MCP layer. 26 tools returning a structured freshness/
completeness envelope, plus high-level tools (
change_impact,find_related_tests,explain_symbol…). → MCP - Investigative visualization. Seeded subgraphs (neighborhood/callers/ callees/impact/domains) with confidence-styled edges and git-diff highlighting — not a decorative hairball. → CLI
Trust is built by being honest about the boundaries:
- ✅ Use it for structural questions — impact, multi-hop call chains, dataflow/taint — and in large or unfamiliar codebases, where reading files by hand is slow and error-prone.
- ✅ Use it when an agent must be sure — a
certainL1 edge is a semantic fact, so the agent can answer and stop instead of re-reading. ⚠️ Reach for grep first when you just want to find a string. For plain text search grep is often enough and cheaper; the graph earns its cost on structure, not substring matching.⚠️ Treat dataflow/taint findings as candidates. It is may-taint — it over-approximates on purpose, so a finding is a lead to verify, not a verdict. The Round 27 Java profile scores 902/0/0/796 on the pinned OWASP matrix and 444/0/0/444 on Juliet CWE-23; all three pinned vulnerable revisions are detected and all three fixes clear. A corrected Juliet project overlay completed with 732/732 files, 4,408certainpromotions and zero resolver warnings/errors. Existing versioned CodeQL rows remain the fair comparison: OWASPdefaultis 776/292/126/504 andsecurity-extendedis 902/471/0/325; Juliet is 222/6/222/438 for both suites. These pinned rows do not establish universal superiority over CodeQL's broader languages, queries, framework models and operational tooling. See Security Benchmark.
Full, quantified limitations and benchmark methodology: FAQ & Limitations and evals/RESULTS.md. What is being worked on next, and why in that order: docs/ROADMAP.md.
Phase-one product languages: Java and Python. The repository also contains 23 dedicated extractors (refined fqn / imports / calls / inheritance): Python, TypeScript/TSX, JavaScript, Rust, Go, Java, Kotlin, C#, C, C++/CUDA/Metal, PHP, Ruby, Lua/Luau, Swift, Scala, Clojure/ClojureScript, Terraform/HCL, and the web tier HTML + CSS/SCSS. These additional languages are experimental until they pass the same end-to-end gates as Java/Python. A generic tier gives structural L0 to dozens more grammars (Zig, Elixir, Vue, Svelte, SQL, Bash, Dart…). Implementation presence is not a claim of product parity. → Languages & Resolvers
| Guide | What's inside |
|---|---|
| Getting Started | Install, index, your first queries |
| Product Contract | Canonical scope, required graph and acceptance gates |
| Core Concepts | The graph model, confidence tiers, the freshness guarantee, layers L0–L3 |
| CLI Reference | Every command, flag, and output format |
| Agents & MCP | The 26 MCP tools and the response envelope |
| Library / Host API | Embedding GraphCodeMap in a service |
| Languages & Resolvers | Language tiers and L1/LSP resolution |
| Semantic Linking Matrix | Real Jedi/JDTLS call contracts and coverage outcomes |
| Product Maturity | Evidence levels, current gaps and honest parity boundaries |
| Language Maturity Playbook | Reusable Java lessons and gates for adapting each next language |
| Round 28 Real-world Feedback | Aethros/PetClinic findings and their Round 29 operational closure |
| Architecture | Pipeline, SQLite schema, incremental indexing |
| FAQ & Limitations | Honest answers, benchmarks, scope |
| Design · Research | Original design contract and research notes |
Alpha (v0.1.0), undergoing a product reset. Java/Python declarations,
parameters, locals, fields/properties and persistent contains/defines/
reads/writes/simple returns now share one structural model. A separate
physical repository graph records every non-ignored folder/file, exact hashes,
index state and Git-aware graph-stage revisions. L1 lifecycle publication is
atomic and observable. Java/Python now persist parameter/local/field → call
argument → callee parameter → return flows_to with atomic publication,
incremental hashes and explicit CFG/heap limitations. Path traversal is the
first vulnerability family on that canonical graph: it supports both explicit
entry parameters and repo-wide persisted source results, with sanitizer-result
cuts and an unknown absence verdict. Migrating the remaining rules,
semantic-link coverage across multiple ordinary
repositories and broad real-world onboarding remain open. Focused real-resolver
matrices currently pass 5/5 Python categories
and 7/7 Java categories (including overload and method-reference resolution).
Flask/PetClinic canaries refine 78.8%/96.9% of persisted local call candidates.
Bounded JDTLS pipelining reduced PetClinic warm revalidation from 54.75s to
37.61s, while unchanged index --l1 runs reuse the snapshot in about 0.10s.
The current focused G3 dogfood (src/codegraph, 1,165 callables) materializes
28,171 value nodes and 37,959 flow edges in 11.446s (historical cold run:
43.04s); five current-snapshot reuses took 0.111–0.119s. Its 80.7% structural path-event
mapping rate is an observability
metric, not a semantic-recall claim.
Historical benchmark results are retained
as bounded subsystem evidence, not as a declaration that the product is ready.
See the Product Contract and
maturity matrix.
Issues and pull requests are welcome — see CONTRIBUTING.md. Wiring a basic language-server adapter is often a ~10-line config. Promoting a language profile requires the structural, semantic, operational and evidence gates in the maturity playbook.
MIT © Victor Alves
Parts of the taint rule catalog are derived from MIT-licensed data published by
other projects (GitHub CodeQL's *.model.yml models and OpenTaint's rules/).
See NOTICE for the required copyright notices and for exactly what
was — and was not — used.