---
title: "AI Memory With Tagging: Tags, Embeddings and Graph Edges Are Not Interchangeable"
url: https://adityaarsharma.com/ai-memory-with-tagging/
date: 2026-10-01
modified: 2026-09-03
lang: en
author: "Aditya Sharma"
description: "MemPalace ships five backends and the default one has a dash in the Namespaces column. Which tools expose tags, and where a tag beats a vector."
categories:
  - "AI"
  - "Automation"
image: https://adityaarsharma.com/wp-content/uploads/2026/09/a2b36eed-95ef-48c5-a806-a66f52281deb_2912x1632-1024x574.webp
word_count: 1934
---

# AI Memory With Tagging: Tags, Embeddings and Graph Edges Are Not Interchangeable

MemPalace ships five storage backends. Its own backend table has a column headed Namespaces. Two of the five have a dash in that column, and one of those two is the default.

`chroma`, the backend you get if you configure nothing, does not support namespaces. Neither does `sqlite_exact`. Milvus, Qdrant and pgvector do. Read from the MemPalace README on 3 September 2026.

![An embedding is a similarity score with no boolean. A tag is a boolean with no similarity.](https://adityaarsharma.com/wp-content/uploads/2026/09/a2b36eed-95ef-48c5-a806-a66f52281deb_2912x1632-scaled.png)That single table is the whole subject of this post. Tagging in AI memory systems is not a feature you turn on.

It is a property of the storage layer underneath, and in most of these projects the tag you think you are writing is either metadata nobody filters on or a scoping key that only some backends honour.

So: three retrieval strategies, what each can and cannot express, which projects expose tags at all, and the specific query shape where a tag beats a vector every time.

On this page

- Why PNG doesn't compress like JPEG does- The Files app "Compress" trap- What actually shrinks a PNG: resize the dimensions- What actually shrinks a PNG: convert it to JPEG- Doing both at once- When you should leave it as PNG- Frequently asked questions
## Three strategies, and the thing each one cannot do
Strip the branding off and there are exactly three ways these systems find the thing you asked for.

| Strategy | What it matches on | What it returns | What it cannot express |
| -------- | ------------------ | --------------- | ---------------------- |
| Tags | Exact set membership | Everything carrying the label, unranked | Similarity. It has no idea that 'auth' and 'login' are related |
| Embeddings | Cosine distance in vector space | The top k nearest, always k of them | Negation, exact sets, or 'none of these are relevant' |
| Graph edges | Traversal from a start node | A path or a subgraph | Anything with no structural connection to your start point |

### The mechanism in one line each
The mechanism is worth stating plainly, because it explains every failure you will hit. An embedding is a similarity score with no boolean. A tag is a boolean with no similarity.

A graph edge is a boolean about a relationship, which is why it can answer 'how are these two connected' and cannot answer 'what did I say about tone of voice'.

### Vector search never says nothing matched
Vector search will always return results. That is not a strength. Ask a vector index for everything tagged as superseded and it will hand you the ten nearest chunks to the word 'superseded', confidently, including nine that are current.

There is no threshold below which it says nothing matched, unless you set one yourself and tune it.

## Which projects expose a tag at all
I went through the documentation for each of these on 3 September 2026 looking for one thing: can I attach a label of my own choosing and then retrieve by that label. The answers vary more than the marketing does.

### Basic Memory: tags in the file, and observations indexed one by one
The most complete tagging model in the field, and it is a text file. Frontmatter carries a `tags` list. Individual observations carry a bracketed category and inline hash tags.

The documented pattern is `- [category] content #optional-tags (optional context)`, and the knowledge format docs state that categories can be anything that makes sense: decision, fact, preference, question, todo, risk, idea, with no fixed list.

`---
tags: [auth, security, backend]
---

## Observations
- [decision] Using JWT tokens for stateless authentication #security
- [constraint] Tokens expire after 15 minutes (based on security audit)`![Basic Memory knowledge format docs showing frontmatter tags and bracketed observation categories](https://adityaarsharma.com/wp-content/uploads/2026/09/36369b08-0c24-4286-b8c0-3012c56a980b_2800x1720-scaled.png)The Basic Memory knowledge format docs, read 3 September 2026. Frontmatter carries a tags list, and each observation carries its own bracketed category.The line in those docs that matters most: each observation is indexed individually.

So the retrieval unit is one bullet, not one file, and each bullet carries its own category and its own tags.

That is a finer grain than any other project here offers, and it is the reason a hand-written taxonomy is workable in this one and painful in the others.

It is also plain text, which means `grep` works, which means your tag query has an escape hatch that does not depend on the vendor's search implementation:

Two lines cover most of it. `grep -rl 'tags:.*superseded' ~/basic-memory/` lists every note whose frontmatter carries that tag, and `grep -rn '^- \[risk\]' ~/basic-memory/` lists every observation filed under risk.

### MemPalace: a scope hierarchy, not free tags
MemPalace does not give you arbitrary labels. It gives you a three-level structure: people and projects become wings, topics become rooms, and original content lives in drawers.

The README's stated purpose for that structure is that searches can be scoped rather than run against a flat corpus, and the CLI exposes it, for example scoping a mining run per project with `--wing`.

![MemPalace README storage backend table with chroma and sqlite_exact showing no namespace support](https://adityaarsharma.com/wp-content/uploads/2026/09/a0e9f499-c20b-4383-bdbb-a028fb76f1f7_2800x1220-scaled.png)The MemPalace README backend table on GitHub, read 3 September 2026. chroma, the default, and sqlite_exact both carry a dash in the Namespaces column.A fixed hierarchy is a real trade against free tags. You cannot label one memory with six orthogonal labels. What you get instead is that the scope is enforced, so it cannot drift the way a free tag vocabulary always does.

And because scoping happens before ranking, a search inside one wing is not competing against your entire history.

The caveat is the backend table at the top of this post. If namespace support is what makes that scoping efficient at your data size, the default embedded ChromaDB is not the configuration you want.

Switching is one environment variable, per the README:

`MEMPALACE_BACKEND=qdrant MEMPALACE_QDRANT_URL=http://localhost:6333 mempalace search "why GraphQL"
# backend and namespace support per the MemPalace README, read 3 September 2026`I compared MemPalace's storage model against mem0's in [the piece on why their benchmark numbers are not measuring the same quantity](https://adityaarsharma.com/mem0-vs-mempalace/), and the short version is that verbatim storage and extracted storage are different products wearing the same word.

### Supermemory and mem0: one scoping key, called something else
Supermemory's README describes memory as scoped with projects, which it calls container tags, so you can separate work and personal context or organise by client or by repository.

In the API that is a `containerTag` passed on both write and read. It is a tag, singular, and it partitions.

mem0's README example passes `filters={"user_id": user_id}` into `memory.search()`. Same idea. A partition key with a filter applied at query time. Neither of these is a taxonomy.

They are tenancy boundaries, and they solve the problem of not mixing two people's memories, which is a different problem from finding the right memory.

### Graphify: tags applied by the system, not by you
Graphify is the interesting case, because it tags things without asking you to. Every edge in its graph carries a provenance marker: `EXTRACTED` if the relationship was explicit in the source, `INFERRED` if Graphify resolved it, `AMBIGUOUS` if it could not decide.

You always know what was read directly and what was guessed.

It also has a second, learned tag layer. `graphify save-result` records how a question and answer turned out, with an outcome of `useful`, `dead_end` or `corrected`.

`graphify reflect --graph` then aggregates those into a lessons document and writes a work-memory overlay, `.graphify_learning.json`, which tags nodes as preferred, tentative or contested, recency-weighted and with provenance.

`graphify save-result --question "Q" --answer "A" --nodes Foo Bar --outcome useful
graphify reflect --graph graphify-out/graph.json
# per the Graphify README, 3 September 2026`That is the most useful tag vocabulary in this whole field and almost nobody talks about it, because a tag that records how a piece of knowledge performed is worth more than a tag that records what topic it is about.

I went through the rest of Graphify's design against its nearest competitor in [the codebase-memory-mcp and Graphify comparison](https://adityaarsharma.com/codebase-memory-mcp-vs-graphify/).

### Graphiti: the tag is time
Graphiti attaches validity windows to facts rather than labels. When information changes, its README states that old facts are invalidated rather than deleted, so you can query what is true now or what was true at a point in time.

Functionally that is a system-applied tag with two values, current and superseded, that you never have to write and cannot forget to apply. The price is the infrastructure: Neo4j, FalkorDB or Neptune, plus an OpenAI key by default.

### codebase-memory-mcp: no tag layer, typed nodes instead
There is no user tagging in codebase-memory-mcp. What it gives you instead is a read-only openCypher subset over typed nodes, so the filter is the node label and the edge type rather than a label you invented:

A supported query reads `MATCH (f:Function)-[:CALLS]->(g) WHERE f.name = 'main' RETURN g.name`, per the codebase-memory-mcp README read on 3 September 2026.

For a code graph that is the right call. The types already exist in the source. Asking a developer to hand-tag functions would be inventing work.

Newsletter

## Automating the boring half

I publish one researched piece a week on putting agents to work on real sites. What I built, what broke, and the commands to check it yourself.

Email address

Get it weekly

Free. One email a week. Unsubscribe in one click, and I do not send anything else.

## Where tagging wins, precisely
A tag beats vector search whenever the set you want is defined by a decision you made rather than by what the text says. That is the rule, and it covers every real case:

- **Provenance.** Everything that came from the client rather than from us. Nothing in the words distinguishes them.- **Status.** Everything superseded. A superseded note and a current note are near-identical in vector space, because they are about the same thing. That is exactly why the embedding cannot separate them.- **Tenancy.** Everything for one client, with a hard guarantee that nothing from another client can rank into the result. A similarity threshold is not a guarantee.- **Confidence.** Everything we are unsure about. This is what Graphify's tentative and contested markers are for.- **Negation.** Everything not in this category. There is no vector for 'not'.
### Where tagging loses
And where it loses, which is most casual use. You have to remember the tag. You have to spell it the same way.

Six months in you have `auth`, `authentication` and `login` as three separate labels covering one idea, and no search finds all three. Vector search has no such failure because it never needed you to be consistent.

A tag vocabulary that nobody prunes is worse than no vocabulary, because it produces confidently incomplete results.

## The shape that works: filter, then rank
The design that survives contact with real data is not a choice between the two. It is an ordering. Apply the boolean first, so the candidate set is correct, then rank within it by similarity, so the order is useful.

### Who already runs filter then rank
Basic Memory's default already does a version of this: hybrid full-text plus vector ranking, with an optional cross-encoder reranking pass over the strongest candidates, both documented in its README and both off by default.

Cognee's architecture is the same principle at a larger scale, combining a graph store with a vector store so structure narrows and embeddings order.

If Cognee's mandatory `LLM_API_KEY` is what stops you using it, the honest replacements are in [the piece on what you are actually replacing when you leave Cognee](https://adityaarsharma.com/cognee-alternatives/).

![Qdrant filtering documentation showing must, should and must_not clauses](https://adityaarsharma.com/wp-content/uploads/2026/09/25fd11b5-600c-430d-9ce1-971e6afcede7_2800x1490-scaled.png)Qdrant's filtering documentation, read 3 September 2026. must, should and must_not are the boolean layer that runs before ranking.
### The failure mode: a tag nothing filters on
What to avoid is a system where the tag exists in the file but nothing filters on it before ranking. That is decoration. You will write tags for a month, trust them, and get results that ignored them.

## One thing to check now
Write two notes into whatever you are using. Identical text, different tag. Then run a search that should return exactly one of them.

If both come back, your tags are metadata rather than a filter, and you should stop writing them or move to a backend that honours them.

That test takes three minutes and it is the only one that settles it.

The wider sort of this field by what each system returns is in [the piece on choosing a memory layer by retrieval unit](https://adityaarsharma.com/best-ai-memory-layer-for-agents/), the storage backends and licences per project are in [the platform comparison table](https://adityaarsharma.com/ai-memory-platform-comparison/), and everything here hangs off [the AI memory tools comparison](https://adityaarsharma.com/ai-memory-tools-compared/).

## Resources
- [Basic Memory knowledge format](https://docs.basicmemory.com/concepts/knowledge-format) - the category and tag pattern, and the per-observation indexing note- [MemPalace on GitHub](https://github.com/MemPalace/mempalace) - the backend table with the Namespaces column- [MemPalace: wings, rooms and drawers](https://mempalaceofficial.com/concepts/the-palace.html) - how scoping is meant to be used- [Graphify on GitHub](https://github.com/Graphify-Labs/graphify) - save-result, reflect, and the preferred / tentative / contested overlay- [Graphiti on GitHub](https://github.com/getzep/graphiti) - validity windows and fact invalidation- [Supermemory on GitHub](https://github.com/supermemoryai/supermemory) - container tags as the scoping primitive- [Qdrant filtering documentation](https://qdrant.tech/documentation/concepts/filtering/) - what a pre-filter over a vector index actually does