BrainBank

code knowledge graph

8/31/2026, 8:58:30 PM · Source

#ai-agents#mcp#use-cases#code-intelligence#codegraph#rust

CodeGraph provides a 100 percent local, pre-indexed semantic code knowledge graph powered by a Rust kernel that drastically cuts token usage and tool calls for AI coding agents.

GitHub - colbymchenry/codegraph: Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, CoPilot, and Hermes Agent — fewer tokens, fewer tool calls, 100% local

GitHub - colbymchenry/codegraph: Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, CoPilot, and Hermes Agent — fewer tokens, fewer tool calls, 100% localGitHub - colbymchenry/codegraph: Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, CoPilot, and Hermes Agent — fewer tokens, fewer tool calls, 100% local

Already installed? Run codegraph upgrade Follow @getcodegraph on X for updates. Supercharge Claude Code, Cursor, Codex, OpenCode, Hermes Agent, Gemini, Antigravity, Kiro, and GitHub Copilot with Semantic Code Intelligence The fastest complete code graph · surgical context · built for how agents actually work · 100% local

  Kernel powered by Rust Documentation & Website →

The CodeGraph platform is coming — for every PR, know exactly what to test, what could break, which flows are affected, and whether business logic is compromised.

Get early beta access to the hosted product · getcodegraph.com

Contents

Get Started Language Support Why CodeGraph? Key Features Read your graph in the browser Framework-aware Routes Mixed iOS / React Native / Expo bridging Quick Start How It Works CLI Reference MCP Tools Library Usage Configuration Telemetry Verified releases Supported Platforms Supported Agents Supported Languages Measured cross-file coverage Troubleshooting License

Get Started

  1. Install the CLI No Node.js required — one command grabs the right build for your OS:

macOS / Linux

Windows (PowerShell)

Already have Node? Use npm instead (works on any version) npm i -g @colbymchenry/codegraph CodeGraph bundles its own runtime — nothing to compile, no native build, works the same everywhere. The installer puts codegraph on your PATH but doesn't change your current shell — open a new terminal before the next step so the command resolves. Upgrade any time with codegraph upgrade — it detects how you installed (bundle, npm, or npx) and updates in place. Add --check to see if an update is available, or codegraph upgrade <version> to pin one.

  1. Wire up your agent(s) In a new terminal, run the installer to connect CodeGraph to the agents you use: codegraph install Detects and auto-configures Claude Code, Cursor, Codex CLI, opencode, Hermes Agent, Gemini CLI, Antigravity IDE, Kiro, and GitHub Copilot (VS Code, Copilot CLI, JetBrains IDEs) — wiring the CodeGraph MCP server into each. This is the step that connects CodeGraph to your agent; installing the CLI in step 1 does not do it on its own. It only wires up your agent — it does not index any code; building each project's graph is the separate codegraph init in step 3. (Shortcut: npx @colbymchenry/codegraph downloads and runs this in one go.)

  2. Initialize each project cd your-project codegraph init codegraph init creates the local .codegraph/ directory and builds the full graph in the same step — one command, done.

  3. No more syncing! Auto-sync is enabled by default. CodeGraph watches the project and updates the graph on every file change — while your agent edits code, or you add, modify, or delete files. The index is never stale, and there is nothing to re-run.

  4. See what your agent sees codegraph ui Opens the graph in your browser at http://127.0.0.1:4747 — callers on the left, the symbol's source in the middle, what it calls on the right. See Read your graph in the browser. Uninstall Changed your mind? One command removes CodeGraph from every agent it configured and the CLI itself — every install it finds (standalone bundle, npm global package, launcher link), shown to you before anything is deleted: codegraph uninstall Pass --keep-cli to remove only the agent configurations and keep the CLI installed. Reverses the installer — strips CodeGraph's MCP server config, instructions, and permissions from each configured agent. Your project indexes (.codegraph/) are left untouched; remove those per-project with codegraph uninit. Use --target to remove from specific agents, or --yes to run non-interactively.

Language Support Every language below gets the same treatment — full structural extraction and cross-file resolution into one graph, no per-language setup:

Per-language details — extensions, frameworks, and what exactly gets extracted — in Supported Languages.

Why CodeGraph? When an AI agent needs to understand code — to answer a question or make a change — it discovers structure the slow way: grep, glob, and Read, one file at a time, rebuilding call paths and dependencies by hand. That's a pile of tool calls and round-trips before it even starts the real work. CodeGraph hands the agent the exact code it needs in one call. It's a pre-built knowledge graph of every symbol, call edge, and dependency in your codebase — so instead of crawling files, the agent asks one question and gets back the relevant source, the call paths between those symbols (including dynamic-dispatch hops grep can't follow), and the blast radius of a change. Surgical context, not a file-by-file search — which means fewer tool calls and faster answers on every codebase, large or small.

A note on cost: CodeGraph's win on every codebase is precision — the agent stops crawling files and answers from the graph. On current models that precision is also a large direct saving: the 2026-08 re-measurement, on a harness that blocks the CLI in both arms, put it at 44% lower cost and 62% fewer tokens on average across the seven benchmark repos, because a strong model without the graph burns its budget re-deriving structure. Cost tracks how much discovery a question demands more than raw repo size: 57–78% on questions the file-reading agent needed 28–43 tool calls to answer, near-even where it got there in 7.

A note on context: the numbers above measure throughput — tokens processed, tools called, dollars spent to reach one answer. They don't measure what is still sitting in your context window afterward, and on that axis CodeGraph costs more, not less. Across the same seven repos in multi-turn sessions, CodeGraph's responses leave about 80% more retrieval context resident at the end of a session than a file-reading agent's do — on VS Code, 67k tokens against 18k. The mechanism is the same one that makes it fast: CodeGraph returns one dense, verbatim payload that answers the question and then stays in the window, where a grep-and-read agent churns through many small results that get evicted. Fewer tokens processed and a larger persistent footprint are both real at once. If you run long sessions in a small window, budget for it. Measured per-repo: docs/benchmarks/residual-context-occupancy.md.

Benchmark Results Tested across 7 real-world open-source codebases spanning 7 languages, comparing an agent (Claude Code, headless) answering one architecture question with and without CodeGraph, at the median of 4 runs per arm. Re-measured 2026-08-05 on Claude Opus 4.8 against the current build, on a harness that blocks the codegraph CLI in both arms — contamination row: 0 of 28 without-arm runs.

The universal win — every repo, every size: 88% fewer tool calls · 53% faster · 62% fewer tokens · 44% cheaper · file reads cut to zero on all seven repos.

With the index available, the agent answers from one to four codegraph_explore calls and stops. Without it, the agent burns its budget on discovery — up to 43 tool calls and 19 file reads re-deriving what the graph already knew. Every repo was faster with CodeGraph in this measurement — by 35% on the narrowest question, by 3.6× on the widest.

Codebase Language Tool calls Time File reads Tokens Cost

VS Code TypeScript · ~11k files 2 vs 28 2.2× faster (58s vs 2m 10s) 0 vs 12 77% fewer 71% cheaper

Excalidraw TypeScript · ~640 2 vs 43 3.6× faster (45s vs 2m 42s) 0 vs 18 84% fewer 78% cheaper

Django Python · ~3k 3 vs 14 35% faster (54s vs 1m 23s) 0 vs 8.5 41% fewer 13% cheaper¹

Tokio Rust · ~790 3 vs 29 2.6× faster (1m 3s vs 2m 43s) 0 vs 19 65% fewer 64% cheaper

OkHttp Java · ~645 1 vs 6 43% faster (33s vs 58s) 0 vs 2 54% fewer 21% cheaper

Gin Go · ~110 1 vs 7 39% faster (28s vs 46s) 0 vs 4 52% fewer ~even¹

Alamofire Swift · ~110 4 vs 33 2.6× faster (54s vs 2m 22s) 0 vs 16.5 59% fewer 57% cheaper

¹ Cost tracks how much discovery the question demanded, which is why it varies far more than the other columns: 57–78% on repos where the file-reading arm needed 28–43 tool calls, but only 13% on Django and even on Gin, where it got there in 14 and 7. The with-arm still answered in 3 and 1 calls with zero file reads. File reads = median files opened — the surgical-context win in one column: the agent never reads a file on any of the seven repos when CodeGraph is present.

Per-repo breakdown — WITH vs WITHOUT (median of 4)

Codebase Metric WITH cg WITHOUT cg

VS Code Time / Tools / Tokens / Cost 58s / 2 / 155k / 0.532m10s/28/670k/0.53 2m 10s / 28 / 670k / 1.80

Excalidraw Time / Tools / Tokens / Cost 45s / 2 / 156k / 0.542m42s/43/991k/0.54 2m 42s / 43 / 991k / 2.43

Django Time / Tools / Tokens / Cost 54s / 3 / 183k / 0.551m23s/14/309k/0.55 1m 23s / 14 / 309k / 0.63

Tokio Time / Tools / Tokens / Cost 1m 3s / 3 / 201k / 0.662m43s/29/573k/0.66 2m 43s / 29 / 573k / 1.83

OkHttp Time / Tools / Tokens / Cost 33s / 1 / 107k / 0.3958s/6/230k/0.39 58s / 6 / 230k / 0.50

Gin Time / Tools / Tokens / Cost 28s / 1 / 87k / 0.3146s/7/180k/0.31 46s / 7 / 180k / 0.31

Alamofire Time / Tools / Tokens / Cost 54s / 4 / 209k / 0.542m22s/33/505k/0.54 2m 22s / 33 / 505k / 1.27

Full benchmark details Methodology. Each arm is claude -p (Claude Opus 4.8, claude-opus-4-8) run headlessly against the repo with --strict-mcp-config: WITH = CodeGraph's MCP server enabled, WITHOUT = an empty MCP config. Built-in Read/Grep/Bash stay available to both. Same question per repo, 4 runs per arm, median reported. Cost = the run's total_cost_usd; Tokens = total tokens processed, summed per assistant turn (input incl. cache reads + cache creation + output); Time = wall-clock; Tool calls = every tool invocation, including those inside any sub-agents the model spawns. Repos cloned at --depth 1 and indexed by the same CodeGraph build that served them. Re-measured 2026-08-05 on the current build. The codegraph CLI is blocked in both arms. A sanitized PATH plus a PreToolUse hook denies any Bash invocation of the CLI, in the WITHOUT arm as well as the WITH arm. This matters: without that block the control arm is not a control. On an unblocked harness we measured the WITHOUT agent finding the CLI on PATH and reaching CodeGraph through Bash in 26 of 28 runs — which distorts the comparison in both directions, since a CLI call is not counted as a tool call and its output still enters the window. Earlier published figures were produced without this block. In the run reported above, all 28 WITHOUT runs attempted the CLI and all 28 were blocked — 0 contaminated. Queries:

Codebase Query

VS Code "How does the extension host communicate with the main process?"

Excalidraw "How does Excalidraw render and update canvas elements?"

Django "How does Django's ORM build and execute a query from a QuerySet?"

Tokio "How does tokio schedule and run async tasks on its runtime?"

OkHttp "How does OkHttp process a request through its interceptor chain?"

Gin "How does gin route requests through its middleware chain?"

Alamofire "How does Alamofire build, send, and validate a request?"

Why CodeGraph wins: with the index available, the agent answers directly — usually one codegraph_explore returns the relevant source — and stops, with zero file reads on every benchmark repo. Without it, the agent spends most of its budget on discovery (find/ls/grep) before reading the right code. CodeGraph only helps when queried directly, so its instructions steer agents to answer directly rather than delegate exploration to file-reading sub-agents — otherwise a sub-agent reads files regardless and CodeGraph becomes overhead.

Built for speed — the Rust kernel CodeGraph's parsing engine is a native Rust kernel: 20 languages — TypeScript, JavaScript, Java, Python, Go, C, C++, Rust, C#, Ruby, PHP, Swift, Kotlin, Scala, Dart, R, Lua, Luau (Metal and CUDA ride the C++ path) — parse in compiled code with one boundary crossing per file. Every language shipped only after its graphs proved byte-for-byte identical to the reference engine on real repositories, from small libraries up to the Linux kernel; platforms without a prebuilt binary and files with syntax errors fall back per-file automatically, same graph either way. And it scales itself to the machine it's on. Worker pools, parallel resolution, and analysis caches are sized from what the system actually has — real core counts (container/cgroup-aware, so a VPS that grants 2 cores gets sized for 2, not the host's 64), honestly-measured available RAM on macOS and Linux, and the measured cost of your project's resolution work:

On a workstation: the full parallel pipeline — native parse workers, a multi-worker resolver pool that engages the moment it pays for itself, memory-gated analysis caches. The Swift compiler repository (27k files of Swift and C++) fresh-indexes in about 100 seconds; a one-file edit re-syncs in ~4. On a 2-core / 6GB VPS: the same graph, from a pipeline tuned to finish — the Linux kernel (70k files, 2M symbols, 6.4M relationships) indexes to completion in under 12 minutes where RAM-first designs run out of memory before reaching 1%. Every day after day one: saving a file updates the graph in well under a second — the watcher fires 300ms after a lone save and syncs exactly what changed (~0.3s of work on a 4,400-file project, ~0.4s on the 27,000-file Swift compiler repo), never re-scanning the tree. Measured against the fastest competing indexer's re-index-on-change: 2–7× faster on medium and larger repos across a 31-repo, 30-language benchmark — and the gap widens with repo size, because their cost grows with the repository and ours grows with the change.

Key Features

Native Rust Kernel Parsing and extraction run in a compiled Rust engine for 20 languages — with graphs verified byte-for-byte identical to the reference engine, and automatic per-file fallback so nothing ever breaks

Adapts to Your Machine Sizes its worker pools and caches from what the system actually has — real core counts (container-aware), honest available RAM, measured per-project cost. A workstation gets the full parallel pipeline; a 2-core VPS gets one tuned to finish reliably

Surgical Context One tool call returns entry points, related symbols, and code snippets — no slow file-by-file exploration

Full-Text Search Find code by name instantly across your entire codebase, powered by FTS5

Impact Analysis Trace callers, callees, and the full impact radius of any symbol before making changes

Always Fresh File watcher uses native OS events (FSEvents/inotify/ReadDirectoryChangesW) with debounced auto-sync — the graph stays current as you code, zero config

20+ Languages TypeScript, JavaScript, ArkTS, Python, Go, Rust, Java, C#, VB.NET, PHP, Ruby, C, C++, CUDA, Objective-C, Metal, Swift, Kotlin, Scala, Dart, Lua, Luau, R, Nix, Erlang, CFML, COBOL, Solidity, Terraform/OpenTofu, Svelte, Vue, Astro, Liquid, Pascal/Delphi

Framework-aware Routes Recognizes web-framework routing files and links URL patterns to their handlers across 17 frameworks

Mixed iOS / React Native / Expo Closes cross-language flows that static parsing misses: Swift ↔ ObjC bridging, React Native legacy bridge + TurboModules + Fabric view components, native → JS event emitters, Expo Modules

100% Local No data leaves your machine. No API keys. No external services. SQLite database only

How auto-syncing works — and why you don't need to run codegraph sync manually When your agent (Claude Code, Cursor, Codex, opencode) launches codegraph serve --mcp, three layers keep the index in step with your code — and make sure the agent never gets a silent wrong answer in the brief window between an edit and the next sync:

File watcher with debounced auto-sync. A native FSEvents / inotify / ReadDirectoryChangesW watcher captures every source-file create / modify / delete and triggers a re-index after a debounce window (default 2000ms, tunable via CODEGRAPH_WATCH_DEBOUNCE_MS, clamped to [100ms, 60s]). Bursts of edits collapse into a single sync.

Per-file staleness banner. During the brief debounce window, MCP tool responses that would reference a still-pending file prepend a ⚠️ banner naming it and telling the agent to Read it directly. Pending files NOT referenced by the response surface as a small footer instead. Either way, the agent gets an explicit signal — validated with Claude Code, where the agent literally says "Reading the file directly for the live content" before opening it.

Connect-time catch-up. When the MCP server (re)connects, codegraph runs a fast (size, mtime) + content-hash reconciliation against the working tree before answering the first query — so edits made while no MCP server was running (a git pull from the terminal, edits from another editor, a previous agent session that exited) get absorbed on the next session's first tool call.

agent writes src/Widget.ts → watcher fires (<100ms) → debounce (default 2s) → sync; Widget.ts is in the index → next agent query sees it

Verify any time with codegraph status (CLI). If anything is pending, you'll see a ### Pending sync: section naming the files and their edit age. The handful of cases where manual codegraph sync makes sense: the watcher is disabled (sandboxed environments, or CODEGRAPH_NO_DAEMON=1), or you're scripting against the index outside an agent session and want a pre-flight sync at the start of your script. → Full deep-dive in Guides → Indexing a Project.

Read your graph in the browser codegraph ui opens a viewer for a project you have already indexed. It is the same graph your agent reads, on screen: pick a symbol and you see who calls it on the left, its verbatim source in the middle, and what it calls on the right — each one drawn level with the line that calls it. codegraph init # once per project, if you haven't already codegraph ui # opens http://127.0.0.1:4747 in your browser

What you get on that screen:

Callers, grouped by file, each with the exact line it calls from — click one to jump there. Test callers fold into a single line so real callers stay in view. The real source, syntax-highlighted, with a marker in the gutter on every line that calls something. Callees on the right, positioned at the line that calls them, joined by a hairline. Hover either end and both light up. Blast radius — direct dependents, everything within three hops, and how many files and test files that touches. Honest edges. A guess CodeGraph isn't sure about is folded away as "uncertain" rather than shown as fact, and a symbol no test reaches within three hops says so. Search (/ or ⌘K) over every symbol and file, and a trail of the path you walked that lives in the URL, so you can send someone the exact route you took. Typing a name also surfaces matching entry points under their own heading, so a URL comes back with the symbol that serves it rather than on its own. Entry points — the first screen on a codebase you have never opened, and the answer to "where does anything start". Every route with its handler and the line it is registered on, grouped by router file and named with the framework it was detected from; the files that run something at import time (a CLI, a worker entry, a script); the tests, ranked by how much of the project each one exercises; and the symbols the most code depends on. Nothing is guessed from a filename — it is all read out of the graph, and a project with no routes says so instead of drawing an empty list. Any row that names a symbol can start a flow: pick a second symbol and you get the path between them, so "how does POST /v1/payroll/cycles/{cycleID}/run reach the database" is two clicks. Click any file path to open the file view: everything that file depends on, its outline in source order, and everything that depends on it. Its Source tab shows the whole file with the same

Learning map

Stage 1: Fundamentals

  • Learn what a code knowledge graph is and how it replaces slow grep-and-read exploration with surgical context.
  • Understand the token and cost savings of providing pre-extracted symbol dependencies directly to AI models.

Stage 2: Installation & Integration

  • Install the CodeGraph CLI on your operating system without native build dependencies.
  • Wire the Model Context Protocol (MCP) server into your preferred AI coding agents like Claude Code or Cursor.

Stage 3: Project Workflows

  • Initialize local indices inside your repositories using project configuration commands.
  • Utilize browser-based graph UI exploration to inspect callers, callees, and blast radii.

Stage 4: Maintenance & Optimization

  • Rely on native OS file watchers and debounced auto-syncing to keep indices fresh automatically.
  • Troubleshoot multi-turn context retention and monitor pending sync statuses.

Get hands-on — step by step

  1. Install the CodeGraph CLI using the curl installer for macOS/Linux or the npm global package manager.
  2. Open a new terminal session and run 'codegraph install' to automatically configure your AI coding agents via MCP.
  3. Navigate into your target project directory and initialize the local knowledge graph using 'codegraph init'.
  4. Launch the local web interface by running 'codegraph ui' and opening http://127.0.0.1:4747 in your browser.
  5. Verify that auto-syncing is functioning properly by checking the status of your project with 'codegraph status'.

Top 3 sources

  1. 1
    CodeGraph GitHub Repository

    The official source code repository, documentation, and release notes for CodeGraph.

    https://github.com/colbymchenry/codegraph

  2. 2
    Model Context Protocol Documentation

    Comprehensive guide on how AI models securely connect to local tools and resources via MCP.

    https://modelcontextprotocol.io/

  3. 3
    Claude Code Documentation

    Official documentation for running Claude Code terminal-based agentic workflows.

    https://docs.anthropic.com/en/docs/agents-and-tools/claude-code

Links are AI-suggested — worth a quick sanity check before diving in.