---
title: "codebase-memory-mcp vs Graphify: What Each One Actually Stores"
url: https://adityaarsharma.com/codebase-memory-mcp-vs-graphify/
date: 2026-09-30
modified: 2026-09-03
lang: en
author: "Aditya Sharma"
description: "Both are alive, both refuse a vector database, and both were pushed to this week. The split is who owns the graph file."
categories:
  - "AI"
  - "Automation"
image: https://adityaarsharma.com/wp-content/uploads/2026/09/79d62277-7c66-4589-8a4f-d2f537207946_2912x1632-1024x574.webp
word_count: 2121
---

# codebase-memory-mcp vs Graphify: What Each One Actually Stores

Two GitHub API calls, four minutes apart on 3 September 2026. `DeusData/codebase-memory-mcp`: 41,957 stars, 3,409 forks, MIT, last release v0.10.8 on 19 August 2026.

`Graphify-Labs/graphify`: 114,056 stars on the first call and 114,057 on the second, 11,086 forks, Apache-2.0, last release v0.9.53 on 30 August 2026. Neither is archived. Neither is abandoned. Both were pushed to within the last four days.

![Neither project is abandoned. That is the answer to the comparison.](https://adityaarsharma.com/wp-content/uploads/2026/09/79d62277-7c66-4589-8a4f-d2f537207946_2912x1632-scaled.png)That is the first finding, and it matters because the query that brought you here assumes one of them is the loser. Neither is.

They are five and six months old respectively, they both build an AST knowledge graph out of your repository, and they both refuse to use a vector database as the primary index.

On the axis people compare them on, they agree.

The difference is what the graph is stored in, who owns the file, and what happens on the second run.

I read both READMEs in full, pulled both repos through the GitHub API, and checked the claims that can be checked from outside. Here is what separates them.

On this page

- The one-line answer- What each one actually stores- The mechanism that decides it: who owns the file- The claims each makes, and which ones survive an outside check- Licence, activity and the things people get wrong- What does not work, in each- Which one to install first- Resources

## The one-line answer
codebase-memory-mcp keeps a per-project SQLite database in a cache directory you never look at, and hands your agent fifteen MCP tools including a read-only Cypher subset.

Graphify writes a `graph.json` file into your working tree and hands you a browsable HTML graph plus a Markdown report. One is a service your agent queries. The other is an artefact you and your agent both read.

If you want code intelligence that is invisible and always current, that points at codebase-memory-mcp.

If you want something that also maps your docs, PDFs and schemas into the same graph and that you can open in a browser, that points at Graphify. The two are only mutually exclusive if you refuse to install both.

## What each one actually stores

### codebase-memory-mcp: SQLite in a cache directory
From the project README, read 3 September 2026: the graph lives in SQLite databases under `~/.cache/codebase-memory-mcp/`, in WAL mode, and persists across restarts.

The override is `CBM_CACHE_DIR`, and the README states that one account can use only one canonical cache root at a time. To reset, you delete the directory.

`# the whole persistence story, from the README
~/.cache/codebase-memory-mcp/ # SQLite, WAL mode, per project
rm -rf ~/.cache/codebase-memory-mcp/ # documented reset`![The GitHub repository page for DeusData/codebase-memory-mcp, MIT licensed and not archived.](https://adityaarsharma.com/wp-content/uploads/2026/09/46c0f306-d4f1-4b5d-8c59-da1219d8cb17_2800x1900-scaled.png)github.com/DeusData/codebase-memory-mcp, screenshot taken 3 September 2026. MIT, not archived.
The indexing pipeline is described as RAM-first: LZ4-compressed reads, an in-memory SQLite database, and a single dump at the end, with memory released back to the operating system afterwards.

Node types are functions, classes, call chains, HTTP routes and cross-service links, parsed with tree-sitter across 162 languages whose grammars are compiled into the binary.

There is a vector layer, and the README is specific about it. The `semantic_query` tool runs vector search across the graph using bundled `nomic-embed-code` embeddings, 768 dimensions at int8, compiled into the executable.

No API key, no Ollama, no Docker. Separately there is BM25 full-text search through SQLite FTS5 with a camelCase and snake_case aware tokenizer. So it is not embedding-free. It is embedding-included.

### Graphify: a JSON file in your repository
Graphify produces three files in `graphify-out/`: `graph.html`, `GRAPH_REPORT.md`, and `graph.json`. The README calls `graph.json` the full graph, and every query command reads from it. There is a documented size cap of 512 MiB, overridable with `GRAPHIFY_MAX_GRAPH_BYTES`.

`uv tool install graphifyy # package graphifyy, command graphify
graphify install # registers the skill with your assistant`
Then run `/graphify .` inside the assistant. Three files land in `graphify-out/`:

- `graph.html`, a clickable force-directed graph.- `GRAPH_REPORT.md`, key concepts and suggested questions.- `graph.json`, the queryable graph itself.![The GitHub repository page for Graphify-Labs/graphify, default branch v8, describing local deterministic AST parsing with no vector store.](https://adityaarsharma.com/wp-content/uploads/2026/09/22ab31fc-9438-4bf4-aabe-5b4cab527153_2800x1900-scaled.png)github.com/Graphify-Labs/graphify, screenshot taken 3 September 2026. Note the default branch is v8 and the description says no vector store.
On vectors, Graphify is blunt in its own README: not a vector index, no embeddings, no vector store, a real graph you traverse. That is a design claim you can check, and the CLI supports it.

`graphify path A B` returns a shortest path with a hop count. `graphify explain X` returns a node with its source file and line, its Leiden community, its degree, and its edges.

The part I find genuinely useful is that every edge carries a provenance tag: `EXTRACTED` if it was explicit in the source, `INFERRED` if Graphify resolved it, `AMBIGUOUS` if it could not decide.

Most tools in this category give you an answer with no marking of how confident the extraction was. This one does.

## The mechanism that decides it: who owns the file
This is where the two designs diverge in a way that shows up in daily use rather than in a feature table.

### Graphify writes into your working tree, so two warnings follow

Because Graphify writes into your working tree, its own README carries two warnings that follow directly from that choice.

First, `graphify hook install` sets up a git merge driver so that `graph.json` union-merges and never shows conflict markers when two developers commit at once. Second, and this one is worth reading twice if you use Claude Code:

> Graphify writes output files into the workspace. If those paths are not ignored, every write invalidates Claude Code's prompt cache, forcing a full re-upload at cache-write rates on the next turn. Add them to .claudeignore. (Paraphrased from the Graphify README, read 3 September 2026.)That is a real cost, documented by the vendor, and it exists only because the graph is a file in your repo. codebase-memory-mcp does not have that failure mode, because nothing it writes is inside your project.

It has a different one: because the state is in a shared cache root, the README describes an admission barrier where all active processes must run the same version, build and cache root, and a genuinely different root is rejected while any process is active.

### The trade, stated plainly

So the trade is legible. File in the repo: visible, diffable, versionable with your code, and it costs you prompt cache.

Database in a cache dir: invisible, no cache churn, and you get version-coupling constraints across every agent session on that machine.

Newsletter

## Automating the boring half

I publish one researched piece a week on putting agents to work on real sites. What I built, what broke, and the commands to check it yourself.

Email address

Get it weekly

Free. One email a week. Unsubscribe in one click, and I do not send anything else.

## The claims each makes, and which ones survive an outside check
Both projects publish headline performance numbers. Both sets are vendor-run. Say that out loud before quoting either.

### The codebase-memory-mcp preprint

codebase-memory-mcp cites a preprint, [Codebase-Memory: Tree-Sitter-Based Knowledge Graphs for LLM Code Exploration via MCP](https://arxiv.org/abs/2603.27277) (arXiv:2603.27277). I checked the arXiv page on 3 September 2026 and the title matches the README exactly.

The claimed result is 83 percent answer quality, 10 times fewer tokens and 2.1 times fewer tool calls than file-by-file exploration, across 31 real-world repositories.

A preprint is not peer review, and the comparison baseline is grep-and-read rather than a rival tool. What I can verify from outside is that the paper exists and says what the README says it says.

### Graphify's benchmark table

Graphify publishes a benchmark table instead: LOCOMO recall@10 of 0.497 against mem0 at 0.048 and supermemory at 0.149, LongMemEval-S QA accuracy of 76 percent, and zero LLM credits to build the graph.

It also states that judging was blind-validated against a second judge with 90.6 percent agreement and a Cohen's kappa of 0.81, which is more methodological disclosure than most projects offer.

It is still the vendor scoring the vendor. I go through what those benchmark names do and do not measure in [the piece on why two systems can report near-identical benchmark scores for different quantities](https://adityaarsharma.com/mem0-vs-mempalace/).

### The zero-credits claim, and where it stops

The zero-credits claim is the one I would treat as the honest headline, because it is structural rather than statistical. Graphify parses code with tree-sitter locally, so the code pass genuinely costs nothing.

Its own capability table says the semantic pass over docs, PDFs, images and video uses your assistant's model or a configured API key.

So the free part is real and the boundary of the free part is documented. Read the boundary, not the badge.

![Bar chart of GitHub stars on 3 September 2026: Graphify at 114,056 and codebase-memory-mcp at 41,957.](https://adityaarsharma.com/wp-content/uploads/2026/09/250cc75d-fe92-47b6-8a88-437730ef0e25_2912x1632-scaled.png)Built from the star counts in this post, pulled from the GitHub REST API on 3 September 2026.

## Licence, activity and the things people get wrong

| | codebase-memory-mcp | Graphify |
| --- | ------------------- | -------- |
| Owner / repo | DeusData/codebase-memory-mcp | Graphify-Labs/graphify |
| Licence | MIT | Apache-2.0 |
| Language | C (single static binary) | Python 3.10+ (PyPI: graphifyy) |
| Repo created | 24 February 2026 | 3 April 2026 |
| Stars, 3 Sep 2026 | 41,957 | 114,056 |
| Forks | 3,409 | 11,086 |
| Open issues | 543 | 1,222 |
| Latest release | v0.10.8, 19 Aug 2026 | v0.9.53, 30 Aug 2026 |
| Last push | 2 September 2026 | 30 August 2026 |
| Archived | No | No |
| Primary store | SQLite in ~/.cache | graph.json in your repo |
| Default branch | main | v8 |
Every figure in that table came from `api.github.com/repos/<owner>/<repo>` and its `/releases` endpoint on 3 September 2026, not from a badge and not from memory. Badges are cached images. Run it yourself:

`curl -s https://api.github.com/repos/DeusData/codebase-memory-mcp \
| python3 -c "import json,sys; d=json.load(sys.stdin); print(d['stargazers_count'], d['pushed_at'], d['archived'])"`Three things people get wrong about this pair, all of which cost me a minute each to check:

- **The Graphify URL you will see quoted is not the canonical one.** Plenty of skill files and blog posts link to `github.com/safishamsi/graphify`. That is a rename. Following the redirect through the API on 3 September 2026 lands on `Graphify-Labs/graphify`. Both work. Only one is the current name.- **Graphify has a v1.0.0 git tag but no v1.0.0 release.** The tags endpoint lists v1.0.0 ahead of v0.9.53, while the releases endpoint's latest is v0.9.53. The README is consistent with this: it advertises early access before a public v1 launch. A tag is not a release.- **codebase-memory-mcp has an active community fork.** `win4r/codebase-memory-mcp-pro`, 223 stars, MIT, last pushed 5 July 2026, describing itself as an incremental-reindex fix plus nine integrated upstream pull requests. It has no published release. If you are on the main project, you are on the maintained one.
## What does not work, in each

### Graphify: Python packaging

Graphify's failure modes are mostly Python packaging, and the README documents them at unusual length.

`uvx graphify install` fails because `uv tool run` reads the first word as a package name and the package is `graphifyy` with two y's. You have to write `uvx --from graphifyy graphify install`.

Separately, if an older `graphifyy` exists in your system Python, `uv run --with graphifyy` can silently load the old copy, which shows up as env overrides being ignored and a 401 that looks like a bad API key.

The fingerprint is a version-mismatch warning line. That is an unusually honest bug to write down.

There is a data-safety behaviour worth knowing too: `graphify extract` refuses to overwrite a larger existing graph with a smaller partial result.

If a walk crashes halfway, your `graph.json` survives. You override with `--allow-partial` or `--force`. That default is the right way round.

### codebase-memory-mcp: antivirus and the version barrier

codebase-memory-mcp's documented rough edge is antivirus. Its own README says Microsoft Defender may flag a release binary as `Trojan:Script/Wacatac.B!ml`, calls it a known false positive, and points to a security document with per-candidate VirusTotal evidence.

Whether you accept that reasoning is your call.

What I will say is that a project that ships a compiled binary, writes to your agent configuration files, and then publishes scan results for every artefact it ships is doing more disclosure than one that says nothing.

The other one is the version barrier. All active processes must run the same exact build and cache root, and an ordinary conflicting process fails before doing work.

If you run Claude Code and Codex side by side and upgrade one, expect to restart the other. That is a documented design decision rather than a bug. It will still surprise you once.

## Which one to install first
If your question is purely about code and you want it to disappear into the background, start with codebase-memory-mcp. One binary, no runtime, a documented reset, and semantic search that needs no key.

If your project's real knowledge is half in Markdown, PDFs and SQL schemas, start with Graphify, because the code pass is free and local and the doc pass is opt-in when you configure a backend.

The browsable `graph.html` is also the fastest way I have found to show a non-engineer what a codebase is shaped like.

### What neither of them is

What neither of them is, is a memory of your conversations. They index artefacts.

If what you want back is the thing you told your agent last Tuesday, that is a different tool class, and I sort them by what they return in [the piece on choosing a memory layer by retrieval unit](https://adityaarsharma.com/best-ai-memory-layer-for-agents/).

The wider field, with storage backend and licence per project, is in [the platform comparison table](https://adityaarsharma.com/ai-memory-platform-comparison/), and the whole cluster hangs off [the AI memory tools comparison](https://adityaarsharma.com/ai-memory-tools-compared/).

One thing to do in the next ten minutes: run the two `curl` commands against the GitHub API for whichever memory tool you are currently trusting, and look at `pushed_at`, `archived`, and the latest release date.

Three fields, thirty seconds, and it settles the maintenance question that every comparison post answers from memory.

## Resources
- [DeusData/codebase-memory-mcp on GitHub](https://github.com/DeusData/codebase-memory-mcp) - README, SECURITY.md, and the release list- [Graphify-Labs/graphify on GitHub](https://github.com/Graphify-Labs/graphify) - README on the v8 default branch- [arXiv:2603.27277](https://arxiv.org/abs/2603.27277) - the codebase-memory-mcp preprint- [graphifyy on PyPI](https://pypi.org/project/graphifyy/) - the package that provides the graphify command- [GitHub REST API: repositories](https://docs.github.com/en/rest/repos/repos) - the endpoint behind every number in this post- [tree-sitter](https://tree-sitter.github.io/tree-sitter/) - the parser generator both projects build on