---
title: "Choosing an AI Memory Layer: Sort Them by What Comes Back, Not by Stars"
url: https://adityaarsharma.com/best-ai-memory-layer-for-agents/
date: 2026-09-30
modified: 2026-09-03
lang: en
author: "Aditya Sharma"
description: "A 24,587-star repo whose README says the code moved elsewhere. One axis that actually separates these tools, and the field sorted on it."
categories:
  - "AI"
  - "Automation"
image: https://adityaarsharma.com/wp-content/uploads/2026/09/910a7c4f-ff18-405a-adfb-40521c46fa33_2912x1632-1024x574.png
word_count: 1811
---

# Choosing an AI Memory Layer: Sort Them by What Comes Back, Not by Stars

`letta-ai/letta` has 24,587 stars. Its README is 1,942 bytes long and it says the repository is now a landing page. The code moved to `letta-ai/letta-code`, which has 3,190 stars.

The last release tagged on the repo with the stars was 0.16.8 on 14 May 2026. The last release on the repo with the code was v0.31.11 on 1 September 2026.

Both figures came from the GitHub API on 3 September 2026.

![The most-starred Letta repository is a 1,942-byte landing page. The code is somewhere else.](https://adityaarsharma.com/wp-content/uploads/2026/09/910a7c4f-ff18-405a-adfb-40521c46fa33_2912x1632-scaled.png)That is the problem with picking a memory layer by looking at a list.

![letta-ai/letta repository page on GitHub showing 24.6k stars and release v0.16.8](https://adityaarsharma.com/wp-content/uploads/2026/09/016969b6-f9ca-450b-824d-021b15453b56_2880x1800-scaled.png)letta-ai/letta on GitHub, screenshot taken 3 September 2026. 24.6k stars, latest release v0.16.8.The list ranks by stars, stars accumulate on the URL people bookmarked in 2024, and the URL people bookmarked is not always where the software is. So this is not a ranking.

![GitHub stars across eight AI memory projects](https://adityaarsharma.com/wp-content/uploads/2026/09/5bb4f667-b118-4edc-8d51-87c577e7c682_2912x1632-scaled.png)GitHub stars across the field, read from the GitHub API on 3 September 2026.It is one axis, applied consistently, with the tools sorted on it.

On this page

- The axis: what comes back when you ask- Sorted on that axis- What running each one actually costs you in attention- The three checks that beat any comparison post, including this one- So which one- Resources
## The axis: what comes back when you ask
Every system in this category claims persistence. All of them persist. That question does not separate anything. The question that separates them is what the retrieval unit is, meaning the shape of the thing that lands back in your context window.

There are four answers in the field right now, and picking between them is the actual decision:

- **A model-written summary of what you said.** Something extracted your sentence into a fact and stored the fact.- **Your original text, unchanged.** The system stored the sentence and gives you the sentence.- **A file you can open in an editor.** The unit is a Markdown note that exists on disk with a name.- **A subgraph or a path.** The unit is a set of nodes and edges, and you get relationships rather than prose.
### Why the four are not interchangeable
These are not interchangeable. If you pick a summariser and later need the exact wording, it is gone.

If you pick a traversal system and ask it a question that has no structural answer, you get nothing useful. Choose the unit first, then the vendor.

## Sorted on that axis
All facts below are from each project's own README or documentation, read on 3 September 2026, with GitHub metadata from the API on the same day.

### 1. Returns a model-written summary
**mem0** (64,601 stars, Apache-2.0, last release ts-v3.1.8 on 2 September 2026). Its README states that mem0 requires a large language model to function, with `gpt-5-mini` as the default, and uses `text-embedding-3-small` as the default embedding model.

### The default vector store
The vector database documentation states that when no configuration is supplied, Qdrant is used.

So the write path is: your text goes to a model, the model produces a memory, the memory is embedded, the embedding goes to Qdrant.

### Supermemory, the same class
**Supermemory** (29,199 stars, MIT, last release server-v0.0.8 on 17 August 2026) sits in the same class.

Its README says it automatically extracts memories and builds user profiles, and explicitly frames that as the point: no embedding pipelines, no vector database configuration, no chunking strategies to choose.

### What you gain, and what you cannot get back
What you gain is compression. A thousand messages become a few dozen facts, and the retrieval is short. What you lose is that the model decided what mattered.

If it drops the caveat in your sentence, the caveat is not recoverable, because the original was never the stored unit. And you pay a model call at write time, every write.

I priced what that arithmetic looks like at volume in [the breakdown of what self-hosted AI memory costs to run](https://adityaarsharma.com/self-hosted-ai-memory-cost/).

### 2. Returns your original text
**MemPalace** (58,808 stars, MIT, last release v3.9.0 on 31 August 2026).

Its README is unusually direct about the mechanism: it stores conversation history as verbatim text and retrieves it with semantic search, and it does not summarise, extract or paraphrase.

### How the index is shaped
The index is structured rather than flat, with people and projects as wings, topics as rooms, and original content in drawers, so a search can be scoped.

The retrieval unit is the drawer, and the drawer holds what you actually wrote.

That is the correct choice when the wording is the information, which is more often than people expect: a decision with its reasoning, a client's exact requirement, a paragraph you liked.

### Where verbatim storage costs you
The cost shows up as volume. Verbatim storage grows with your history rather than with the number of distinct facts you hold.

The default retrieval backend is ChromaDB, embedded and local, with a bundled `sqlite_exact` alternative, and Milvus, Qdrant and pgvector available as opt-in server backends.

Embeddings run locally by default, roughly 30 MB for the English-only MiniLM model or about 300 MB for the multilingual EmbeddingGemma option, per the README.

### 3. Returns a file you can open
**Basic Memory** (3,841 stars, AGPL-3.0, last release v0.23.2 on 25 August 2026). The smallest star count on this page and the clearest storage model on it.

### Files on disk plus a SQLite index
Its README describes the store as plain Markdown files on your disk plus a local SQLite index, with no servers required.

Each file is an entity with observations and relations, and the relations are wiki-style double-bracket links.

`---
title: Coffee Brewing Methods
permalink: coffee-brewing-methods
tags: [coffee, brewing]
---

## Observations
- [method] Pour over gives the cleanest cup

## Relations
- pairs_with [[Grinder Settings]]`
### You can fix a wrong memory by editing a file
The reason this matters more than it sounds: when the retrieval unit is a file,

you can fix a wrong memory by editing a file.

In every summariser, correcting a bad memory means asking the system to correct it and hoping. Here you open the note in Obsidian and change the line.

The README says Obsidian needs no setup at all: point it at the project folder and the same wikilinks and frontmatter appear in its graph view.

### The trade
The trade is that you are now maintaining a knowledge base. Files accumulate, links go stale, and nobody prunes them. A summariser at least has an opinion about what to keep.

### 4. Returns a subgraph or a path
**Graphiti** (30,536 stars, Apache-2.0, last release mcp-v1.1.0 on 1 September 2026) stores facts as triplets with temporal validity windows, plus the raw episodes that produced them.

### Facts that expire
Its documented behaviour when information changes is that old facts are invalidated rather than deleted, so you can query what is true now or what was true at a point in time.

That is the only system here whose retrieval unit carries its own expiry.

### The database you have to keep running
The operational cost is real and stated in the README: it needs a graph database.

Neo4j 5.26, FalkorDB 1.1.2, Amazon Neptune with OpenSearch Serverless, or Kuzu 0.11.2 which the README marks deprecated because the upstream project is no longer maintained.

It also defaults to OpenAI for both inference and embeddings and expects an API key.

### Cognee, and the API key requirement
**Cognee** (30,426 stars, Apache-2.0, last release v1.5.3 on 23 August 2026) is the same shape with lighter defaults: the documentation lists Ladybug/Kuzu as the default graph store and LanceDB as the default vector store, with SQLite relational, all embedded for local development.

It requires `LLM_API_KEY` to index. If that constraint is what is pushing you away from it, I went through the actual replacements in [the piece on what you are actually replacing when you leave Cognee](https://adityaarsharma.com/cognee-alternatives/).

**LightRAG** (39,349 stars, MIT, last release v1.5.7 on 2 September 2026) and the two code-graph tools, codebase-memory-mcp and Graphify, also sit here. The code-graph pair is a different job entirely, and I compared them directly in [the codebase-memory-mcp and Graphify comparison](https://adityaarsharma.com/codebase-memory-mcp-vs-graphify/).

Newsletter

## Automating the boring half

I publish one researched piece a week on putting agents to work on real sites. What I built, what broke, and the commands to check it yourself.

Email address

Get it weekly

Free. One email a week. Unsubscribe in one click, and I do not send anything else.

## What running each one actually costs you in attention
Price is the wrong first question, because at personal scale all of these are cheap and at production scale the model calls dominate everything else. The better question is how many moving parts you have agreed to keep alive.

| Retrieval unit | Example | Services you must run | Model call at write time |
| -------------- | ------- | --------------------- | ------------------------ |
| Model summary | mem0 | A vector store (Qdrant by default) | Yes, gpt-5-mini by default |
| Verbatim text | MemPalace | None. ChromaDB is embedded | No. Local embeddings by default |
| Markdown file | Basic Memory | None. Files plus SQLite | No for writing. Optional for semantic search |
| Graph triplet | Graphiti | Neo4j, FalkorDB or Neptune | Yes, OpenAI by default |
| Graph triplet | Cognee | None locally. Kuzu and LanceDB embedded | Yes, LLM_API_KEY required |

### Reading the last two columns
![1,942 bytes of README on a 24,587 star repository](https://adityaarsharma.com/wp-content/uploads/2026/09/1c53254c-7f87-4cb4-8b7f-bb5566817a42_2400x2400.png)The README that told me Letta had moved. GitHub API, 3 September 2026.Read the last two columns together.

The systems with no service to run and no model call at write time are the ones you will still be using in six months, because there is nothing to forget to restart and no bill that grows while you sleep.

That is not an argument that they retrieve better. It is an argument that the failure mode of the others is abandonment.

## The three checks that beat any comparison post, including this one
Every list in this category, mine included, is a snapshot. Here is how to re-derive it in under a minute per project.

`R=getzep/graphiti # or whichever
curl -s https://api.github.com/repos/$R | \
python3 -c "import json,sys;d=json.load(sys.stdin);print(d['pushed_at'],d['archived'],(d['license'] or {}).get('spdx_id'))"
curl -s https://api.github.com/repos/$R/releases/latest | \
python3 -c "import json,sys;d=json.load(sys.stdin);print(d.get('tag_name'),d.get('published_at'))"`Three things to read off that output.

- **Compare `pushed_at` with the latest release date.** A repo pushed yesterday with no release in five months is being worked on but not shipped. Khoj is the clean example: pushed 2 August 2026, latest published release 2.0.0-beta.28 dated 26 March 2026.- **Read the README's own length.** A 1,942-byte README on a 24,587-star project told me Letta had moved before any article did.- **Search the README for the word 'key'.** If an API key is required to index, the system has an ongoing cost and a network dependency whatever else it claims about being local.
## So which one
If you want your own words back, MemPalace.

If you want to be able to open and correct what the agent remembered, Basic Memory.

If you need to answer questions about when something stopped being true, Graphiti, and accept that you are now running a graph database.

If you want the shortest possible retrieval and do not mind a model deciding what mattered, mem0.

If you are still deciding whether you need a memory layer at all rather than retrieval over documents, start one step earlier with [the difference between an MCP server and a RAG pipeline](https://adityaarsharma.com/mcp-server-vs-rag/), because a good half of the traffic on this topic is asking a protocol question in architecture words.

The full field with storage backends and licences is in [the platform comparison table](https://adityaarsharma.com/ai-memory-platform-comparison/), and the cluster hub is [the AI memory tools comparison](https://adityaarsharma.com/ai-memory-tools-compared/).

One thing to do now: open whichever memory tool you already have installed, save a sentence with an unusual phrase in it, then search for it and see whether the sentence comes back or a summary of the sentence comes back.

That single test tells you which of the four classes you are in, and it takes about two minutes.

## Resources
- [mem0](https://github.com/mem0ai/mem0) - README and the default model configuration- [mem0 vector store documentation](https://docs.mem0.ai/components/vectordbs/overview) - the line stating Qdrant is the default- [MemPalace](https://github.com/MemPalace/mempalace) - storage backend table and verbatim storage claim- [Basic Memory](https://github.com/basicmachines-co/basic-memory) - the Markdown entity format- [Graphiti](https://github.com/getzep/graphiti) - graph database requirements and temporal fact handling- [Cognee](https://github.com/topoteretes/cognee) - default Kuzu and LanceDB backends- [letta-ai/letta-code](https://github.com/letta-ai/letta-code) - where the Letta source actually lives now- [GitHub REST API: releases](https://docs.github.com/en/rest/releases/releases) - for the release-date check above