
MemPalace
話された内容をそのまま保存し、要約するのではなく索引の側を構造化する、手元で動く記憶の仕組み
MemPalaceとは
多くの記憶の仕組みは、何を覚える価値があるかをエージェントに判断させ、その要約を保存します。つまり、重要だと思われなかった細部は、必要になったときにはもう残っていません。MemPalaceは逆の道を選びます。会話は一言一句そのまま保存し、構造は内容ではなく索引の側に持たせます。人やプロジェクトが棟に、話題が部屋になり、元の文章は引き出しに収まります。そのため検索を全体に平らにかけるのではなく、書庫の一部に向けて絞り込めます。動作は手元のマシンで完結し、MCPとして公開できるためコーディングエージェントから直接検索できます。保存先はChroma、SQLite、Milvus、Qdrant、Postgresから選べます。取り込みは自動で起こるものではなく、自分で実行する作業です。
MemPalaceで何ができますか?
- 元の言い回しをそのまま残す — 取り込む段階で要約も言い換えもしないため、当時は重要に見えなかった細部が、必要になったときにも残っています。
- 書庫の一部に絞って検索する — 索引は人、プロジェクト、話題で整理されているため、これまで保存したすべてと競わせずに範囲を限って検索できます。
- コーディングエージェントから直接検索する — MCPサーバーとして動くため、Claude Codeなどのクライアントが作業の途中で、別の手順や貼り付けを挟まずに参照できます。
- 過去のセッションをディスクから取り込む — プロジェクトのファイルや既存の会話記録を取り込めるので、記憶は空の状態からではなく、これまでの経緯から始まります。
- 保存先を選ぶ — 既定は設定不要の組み込みデータベースです。1台では収まらなくなった場合に備えて、Milvus、Qdrant、Postgresにも対応しています。
- 手元のマシンから出さない — 索引作成も検索も、取得した埋め込みモデルを使って手元で動きます。自分で選ばない限り、どこにも送信されません。
MemPalaceを選ぶ前に
- そのまま保存する方式なので、取り込んだ分だけ書庫は大きくなります。取り込みは自動ではなく自分で実行するコマンドのため、内容は最後に更新した時点までしか新しくありません。
- 開発元は、名前を似せた別ドメインが本物を装い、不正なソフトウェアを配布する恐れがあると注意しています。正規の入手先はリポジトリ、PyPI、公式の資料サイトのみと明示されています。
よくある質問
MemPalaceは商用利用できますか?
MemPalaceはMITライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。
MemPalaceはどの形で使えますか?
MemPalaceはローカル実行・セルフホストの形で利用できます。
ドキュメント
MemPalace/mempalace のREADMEより転載(MIT)。 原文を読む ↗
MemPalace
Local-first AI memory. Verbatim storage, pluggable backend, 96.6% R@5 raw on LongMemEval — zero API calls.
[![][version-shield]][release-link] [![][python-shield]][python-link] [![][license-shield]][license-link] [![][discord-shield]][discord-link]
[!CAUTION] Beware of impostor sites. MemPalace has no other official websites. The only official sources are this GitHub repository, the PyPI package, and the docs at mempalaceofficial.com. Any other domain (including
.tech,.net, or other.comvariants) is an impostor and may distribute malware. Details and timeline: docs/HISTORY.md.
[!IMPORTANT] Claude Code sessions expire in 30 days without auto-save hooks wired. Read this →
Need the shortest recovery/setup path? Use the Claude Code retention setup checklist.
What it is
MemPalace stores your conversation history as verbatim text and retrieves it with semantic search. It does not summarize, extract, or paraphrase. The index is structured — people and projects become wings, topics become rooms, and original content lives in drawers — so searches can be scoped rather than run against a flat corpus.
The retrieval layer is pluggable. The current default is ChromaDB; the
interface is defined in mempalace/backends/base.py
and alternative backends can be dropped in without touching the rest of
the system.
Nothing leaves your machine unless you opt in.
Architecture, concepts, and mining flows: mempalaceofficial.com/concepts/the-palace.
Install
MemPalace ships a CLI, so install it in an isolated environment to avoid
PEP 668 errors on Debian/Ubuntu/Homebrew Pythons and to keep mempalace’s
deps (chromadb, numpy, grpcio, …) from conflicting with anything
else in your global site-packages.
We recommend uv — uv tool install puts
the mempalace CLI in an isolated environment on your PATH:
uv tool install mempalace
mempalace init ~/projects/myapp
pipx works the same way if you prefer it:
pipx install mempalace.
Prefer plain pip only inside an activated virtualenv where you
explicitly want import mempalace available:
python -m venv .venv && source .venv/bin/activate
pip install mempalace
Android / Termux
Native Termux installation is not currently supported because compiled dependencies such as ChromaDB and ONNX Runtime publish Linux wheels, not Android wheels. Android ARM64 users can run the regular Linux packages in an isolated Debian PRoot container instead. See the Termux installation guide for the tested setup and an argv-preserving launcher.
Docker
A container image is also available for running the MCP server or the CLI without a local Python toolchain. Multi-arch (amd64 + arm64), so it runs natively on Apple Silicon:
docker pull ghcr.io/mempalace/mempalace:latest
Everything persists under /data — palace, config, and the cached embedding
model — so mount a volume there and reuse it across runs:
# MCP server over stdio — note the `-i` flag (JSON-RPC needs stdin)
docker run -i --rm -v mempalace-data:/data ghcr.io/mempalace/mempalace
# Run any CLI command instead. The container only sees what you mount, so
# mount the directory you want to mine — read-only is enough, mining never
# writes to the source.
docker run --rm -v mempalace-data:/data -v /path/to/project:/work:ro \
ghcr.io/mempalace/mempalace mine /work
docker run --rm -v mempalace-data:/data ghcr.io/mempalace/mempalace search "why GraphQL"
The first command that needs embeddings downloads the model into /data
(~80 MB for the default minilm, ~300 MB for embeddinggemma). It is a
one-off as long as the volume persists, but it does mean the first call is
slow and needs network — worth knowing before assuming a hung container.
Wire it into an MCP client (e.g. Claude Code) as a stdio server. Mount anything you want the server to be able to mine — it cannot reach your transcripts otherwise:
{
"mcpServers": {
"mempalace": {
"command": "docker",
"args": [
"run", "-i", "--rm",
"-v", "mempalace-data:/data",
"-v", "/absolute/path/to/.claude/projects:/transcripts:ro",
"ghcr.io/mempalace/mempalace"
]
}
}
}
Use a real absolute path there — ~ and $HOME are not expanded by every
MCP client. Paths are container paths from then on: mine /transcripts, not
~/.claude/projects.
Mount permissions on Linux. The image runs as uid 1000 and bind mounts
keep their host ownership, so a mounted directory has to be readable by that
uid — an ordinary 0755 checkout is fine, a 0700 directory is not, and the
failure surfaces as PermissionError: [Errno 13] rather than anything about
Docker. Docker Desktop maps uids on macOS and Windows, so this only bites on
Linux. Do not work around it with --user: /data is owned by uid 1000
inside the image, so another uid cannot write the palace at all.
docker compose run --rm mcp works too (see docker-compose.yml), and
deploy/docker-compose.server.yml stands up the team server. To build the
image yourself instead of pulling — required for the GPU variant, which is not
published:
docker build -t mempalace . # CPU
docker build --build-arg EXTRAS="extract,spellcheck" -t mempalace .
docker build -f Dockerfile.gpu -t mempalace:gpu . # CUDA; run with --gpus all
The GPU image is x86_64-only: onnxruntime-gpu publishes no aarch64 Linux
wheels, so that last build fails on an ARM host (including Apple Silicon) with
a dependency-resolution error rather than an obvious one.
Note that a build from a clone uses whatever branch you checked out; develop
is the default branch, so pull the published image if you want the released
version.
Storage backends
ChromaDB is the default and needs no configuration. MemPalace also ships a pluggable backend contract, exercised across deliberately different substrates so the contract is never accidentally shaped around one vendor. Every non-default backend is opt-in.
| Backend | Mode | Install | Namespaces | Lexical | Configure with |
|---|---|---|---|---|---|
chroma (default) | Local (embedded) | bundled | – | ✓ | – |
sqlite_exact | Local (exact) | bundled | – | ✓ | – |
milvus | Local (Lite) · Server opt-in | mempalace[milvus] | ✓ | ✓ | MEMPALACE_MILVUS_URI |
qdrant | Server (REST) | bundled | ✓ | ✓ | MEMPALACE_QDRANT_URL |
pgvector | Server (Postgres) | mempalace[pgvector] | ✓ | ✓ | MEMPALACE_PGVECTOR_DSN |
Select with --backend <name>, MEMPALACE_BACKEND=<name>, or
"backend": "<name>" in config.json. See
Storage backends for connection
variables, namespace behavior, and deployment notes.
Quickstart
# Mine content into the palace
mempalace mine ~/projects/myapp # project files
mempalace mine ~/.claude/projects/ --mode convos # Claude Code sessions (scope with --wing per project)
# Search
mempalace search "why did we switch to GraphQL"
# Load context for a new session
mempalace wake-up
For Claude Code, Gemini CLI, Antigravity, MCP-compatible tools, and local models, see mempalaceofficial.com/guide/getting-started.
Benchmarks
All numbers below are reproducible from this repository with the commands
in benchmarks/BENCHMARKS.md. Full
per-question result files are committed under benchmarks/results_*.
LongMemEval — retrieval recall (R@5, 500 questions):
| Mode | R@5 | LLM required |
|---|---|---|
| Raw (semantic search, no heuristics, no LLM) | 96.6% | None |
| Hybrid v4, held-out 450q (tuned on 50 dev, not seen during training) | 98.4% | None |
| Hybrid v4 + LLM rerank (full 500) | ≥99% | Any capable model |
The raw 96.6% requires no API key, no cloud, and no LLM at any stage. The hybrid pipeline adds keyword boosting, temporal-proximity boosting, and preference-pattern extraction; the held-out 98.4% is the honest generalisable figure.
The rerank pipeline promotes the best candidate out of the top-20
retrieved sessions using an LLM reader. It works with any reasonably
capable model — we have reproduced it with Claude Haiku, Claude Sonnet,
and minimax-m2.7 via Ollama Cloud (no Anthropic dependency). The gap
between raw and reranked is model-agnostic; we do not headline a “100%”
number because the last 0.6% was reached by inspecting specific wrong
answers, which benchmarks/BENCHMARKS.md flags as teaching to the test.
Other benchmarks (full results in benchmarks/BENCHMARKS.md):
| Benchmark | Metric | Score | Notes |
|---|---|---|---|
| LoCoMo (session, top-10, no rerank) | R@10 | 60.3% | 1,986 questions |
| LoCoMo (hybrid v5, top-10, no rerank) | R@10 | 88.9% | Same set |
| ConvoMem (all categories, 250 items) | Avg recall | 92.9% | 50 per category |
| MemBench (ACL 2025, 8,500 items) | R@5 | 80.3% | All categories |
We deliberately do not include a side-by-side comparison against Mem0, Mastra, Hindsight, Supermemory, or Zep. Those projects publish different metrics on different splits, and placing retrieval recall next to end-to-end QA accuracy is not an honest comparison. See each project’s own research page for their published numbers.
Reproducing every result:
git clone https://github.com/MemPalace/mempalace.git
cd mempalace
uv sync --extra dev # or: pip install -e ".[dev]"
# see benchmarks/README.md for dataset download commands
uv run python benchmarks/longmemeval_bench.py /path/to/longmemeval_s_cleaned.json
Knowledge graph
MemPalace includes a temporal entity-relationship graph with validity windows — add, query, invalidate, timeline — backed by local SQLite. Usage and tool reference: mempalaceofficial.com/concepts/knowledge-graph.
MCP server
44 MCP tools cover palace reads/writes, knowledge-graph operations, cross-wing navigation, drawer management, agent diaries, and agent coordination (logstream events + artifact handoffs). Installation and the full tool list: mempalaceofficial.com/reference/mcp-tools.
Agents
Each specialist agent gets its own wing and diary in the palace.
Discoverable at runtime via mempalace_list_agents — no bloat in your
system prompt:
mempalaceofficial.com/concepts/agents.
Auto-save hooks
Auto-save hooks for Claude Code, Codex CLI, and Cursor IDE save periodically and before context compression:
- Claude Code + Codex → mempalaceofficial.com/guide/hooks
- Cursor IDE (adds session-start recall and a transcript snapshot before compaction) → mempalaceofficial.com/guide/cursor-hooks
If you are installing under time pressure, start with the
Claude Code retention setup checklist:
wire the hooks, back up existing JSONL transcripts, and backfill them with
mempalace mine ~/.claude/projects/ --mode convos.
For per-message recall on top of the file-level chunks the hooks produce,
run mempalace sweep <transcript-dir> periodically — it stores one
verbatim drawer per user/assistant message, idempotent and resume-safe.
Requirements
- Python 3.9+
- A vector-store backend (ChromaDB by default)
- ~300 MB disk for the embedding model. Onboarding (
python -m mempalace.onboarding) offersembeddinggemma-300m(multilingual, 100+ languages, recommended) orall-MiniLM-L6-v2(English-only, ~30 MB). See the docstring atmempalace/embedding.pyfor details and migration notes. - Optional — compute embeddings on a server instead of locally. Set
embedding_model: "openai-compat"in~/.mempalace/config.jsontogether withembedding_api_url/embedding_api_model(andembedding_api_keyif the server needs auth) to use any OpenAI-compatible/v1/embeddingsendpoint — LM Studio, llama.cpp, vLLM, Ollama’s OpenAI shim, or a self-hosted server (e.g. a larger multilingual or GPU-served embedder). Each key is overridable via the matchingMEMPALACE_EMBEDDING_API_*env var. When the endpoint is on your machine or LAN, no content leaves your network. Switching to it requiresmempalace repair rebuild-index(different vector space).
No API key is required for the core benchmark path.
Docs
- Getting started → mempalaceofficial.com/guide/getting-started
- CLI reference → mempalaceofficial.com/reference/cli
- Python API → mempalaceofficial.com/reference/python-api
- Full benchmark methodology → benchmarks/BENCHMARKS.md
- Release notes → CHANGELOG.md
- Corrections and public notices → docs/HISTORY.md