Introduction
Overview
ContextGraph is a local-first code understanding engine for AI agents. It parses your repository into a graph of symbols and the resolved edges between them — built from tree-sitter ASTs, with scope-correct identity, so a getName in one class is never confused with a getName in another — and stores it in SQLite on your disk.
Agents reach it through a Model Context Protocol (MCP) server. The primary tool, contextgraph.explore, takes a natural-language question and returns the matched modules, the relevant symbols with verbatim source, the resolved edges between them, and the blast radius of changing them — in one call.
The same engine also ingests Markdown, PDFs, SQL schemas and config files into the same graph, so a design document and the code implementing it end up connected rather than merely stored side by side.
Why it exists
Large codebases are hard to navigate — for humans and AI agents alike. Reading every file to answer "what does this module do?" or "what calls this function?" is slow and expensive, and grep answers with text matches rather than with meaning.
Most graph tools then hand an agent a list of pointers and let it go read files anyway, which spends the tokens the graph was supposed to save. ContextGraph returns the source itself, ranked against the question that was asked, capped at a token budget. Across 33 questions over four repositories, all four sides measured in one run, its mean reciprocal rank over the 29 headline questions is 0.502, against 0.326 for CodeGraph (third-party), 0.218 for a stock grep and 0.218 for ripgrep (third-party). It does not win everywhere: on gin CodeGraph (third-party) leads on MRR, 0.813 against 0.719; on keycloak bash (base-system shell only) and ripgrep (third-party) lead on recall@10, 39.6% against 17.3%. The project README spells out what the numbers do and do not establish.
What's in the graph
Every indexed file produces nodes and edges:
- Artifact nodes — the file itself (
CodeFile,MarkdownFile,PDF,DatabaseSchema,ConfigFile…) - Entity nodes — symbols extracted from the file (
Class,Function,Method,DatabaseTable,Concept,Claim…) - Edges — relationships between nodes (
DEFINES,CALLS,DEPENDS_ON,REFERENCES,CONTAINS…)
Every node carries a confidence score and full provenance — the exact file path and line numbers where it was extracted from. Every edge additionally records the resolution rung that produced it: references that cannot be resolved within a single file are resolved afterwards through a ladder of increasingly permissive strategies, and an edge tells you which one matched. That is what lets you distinguish a call resolved by exact type from one resolved by name alone.
Extractors
| Extractor | File types | What it extracts |
|---|---|---|
TreeSitterExtractor | Kotlin, Java, TypeScript, TSX, JavaScript, Python, Swift, Objective-C, Go | Classes, functions, methods, imports, call sites — with scope-correct symbol identity |
MarkdownExtractor | .md | Headings, sections, links |
SqlExtractor | .sql | Tables, columns, foreign keys |
PdfExtractor | Sections, text content | |
ConfigExtractor | .yaml, .yml, .json, .toml | Keys, values, structure |
SemanticExtractor | any (via LiteLLM) | Concepts, claims, decisions, relationships |
100% local
ContextGraph runs entirely on your machine. The graph is stored as SQLite files under .contextgraph/ in your project directory — graph.local.db, your working copy, and optionally a committed graph.db baseline that CI writes so a fresh clone is queryable before anyone indexes anything. No data is sent to any cloud service unless you explicitly enable the optional semantic extractor and point it at your own LiteLLM endpoint.