claude-mem: A Compression Layer for Persistent Cross-Session Agent Memory

Published · AI Daily — AI-assisted deep research, methodology & disclosure

claude-mem records an agent's tool-use observations during a session, compresses them into semantic summaries with AI, and injects only the relevant part when the next session starts, so the agent no longer begins from zero. The project has about 96,000 stars and supports hosts such as Claude Code, OpenClaw, Codex, Gemini and OpenCode. Compression can run on its own observer, an OpenRouter or Gemini key, or an Anthropic plan. The README gives no reproducible evaluation, so test it on your own repository and audit where sensitive data is stored and sent.

Background and Problem Definition

Coding agents share one structural weakness: when a session ends, the context is gone. A developer spends an hour teaching an agent the layout of a repository, the traps in its build, and the trade-offs behind a design. The next morning the agent starts from zero. The usual fixes all have a cost. A hand-written instruction file such as CLAUDE.md depends on a person to keep it current, and it drifts. Replaying the full chat history into a new session fills the context window fast and burns tokens. claude-mem targets this gap. It keeps what an agent did across sessions and injects only the relevant part into the next one.

The project is maintained by thedotmack and has about 96,000 GitHub stars. Its topics include ai-memory, long-term-memory, chromadb, sqlite, rag and claude-code-plugin. According to the README, it began as a persistent memory compression system for Claude Code. It now lists OpenClaw, Codex, Gemini, Hermes, Copilot and OpenCode as hosts, with install targets for Grok Bot, Antigravity CLI and OMP as well. One caveat up front: this article rests on the repository README and metadata. The author did not deploy the tool or run a benchmark, so every statement about internals below is limited to what the README says.

Architectural Core and Technical Principles

The README describes a three-step loop: capture, compress, inject. First, the system records observations of tool usage during a session. Second, it uses AI to compress those raw records into semantic summaries. Third, when a new session starts, it makes the relevant compressed context available to the agent. The repository topics name sqlite and chromadb. That suggests structured records in SQLite and semantic retrieval in a vector store, a common hybrid design. The table layout and the ranking strategy cannot be confirmed from the README excerpt, so treat that part as inference.

The integration model is the second design point. For hosts that expose hooks, such as Claude Code, the installer registers plugin hooks and starts a worker service. That service handles observation and compression in the background, so the main session is not blocked. For hosts without hooks, the README says the tool watches chat log files instead. Grok Bot is the named example. This is a pragmatic adaptation. It needs no event API from the host. The price is that it reads the record after the fact, so freshness and structure depend on the log format.

The third point is who does the compression. The README lists several memory providers: the project's own claude-mem observer, your own OpenRouter or Gemini key, or your Anthropic plan. A local observer is opt-in through the --provider host flag. This puts a hidden cost on the table. Compressing memory means calling a language model, and that call is paid for either with your plan quota or with money at a third party.

Practical Evaluation and Applications

The install path is short. The general command is npx claude-mem install. You can add --ide opencode, --ide antigravity, --ide omp or --ide grok-bot to pick a host. Inside Claude Code, run /plugin marketplace add thedotmack/claude-mem, then /plugin install claude-mem, and restart. The README warns about one trap in particular: npm install -g claude-mem installs only the SDK and library. It does not register hooks and does not start the worker service. You must use the npx installer or the plugin commands. The account flow deserves attention too. By default, the installer finishes setup and then asks you to sign in through the browser with an email magic link. Signing in unlocks a 14-day free trial of the claude-mem observer. When the trial ends, memory falls back to your Anthropic plan unless you subscribe. To skip sign-in, pass an explicit --provider flag, set CLAUDE_MEM_ONLINE_OPTIN=false, or run in CI or a non-interactive shell. A team with privacy or compliance duties should make that choice on purpose at deployment time, not accept the default.

The README also describes an awareness push pilot. Observations judged important, with types decision, bugfix, security_alert and sensitive, are appended as dated lines to a monthly log for Grok Bot. The feature can be turned off with CLAUDE_MEM_GROK_BOT_AWARENESS_ENABLED=false. Note the sensitive category: content of that type lands in a log file, so users should check redaction and file permissions themselves. The honest conclusion is that the README gives no reproducible comparison. We do not know how often compression drops a key detail, how often it hardens a wrong conclusion, or how good retrieval is. Before adoption, run a small trial on your own repository and compare task quality and token use with and without memory.

Industry Impact and Outlook

A star count near 100,000 shows that cross-session memory is among the strongest needs in the agent toolchain. Topic tags such as mem0, supermemory and openmemory show a crowded field. The differentiator for claude-mem is breadth of hosts. It places one memory layer under many agents instead of tying itself to one product. That appeals to teams that use several coding agents, because experience can move between tools.

The risks are also clear. First, a memory layer becomes a new trust boundary. It reads every tool result, and those may hold secrets and private code. Second, hosted defaults raise a data-flow question: users must confirm which parts pass through a third-party service. Third, compression is lossy, and a wrong summary gets injected again and again as if it were fact. Likely next steps are auditable memory entries, memory that a user can edit and delete, and public benchmarks. Until those exist, treat claude-mem as a tool worth trying and worth auditing, not as mature infrastructure.

Sources

FAQ

Why does npm install -g claude-mem not enable the memory feature?

Per the README, that command installs only the SDK or library. It does not register plugin hooks and does not start the worker service. Use npx claude-mem install, or run /plugin marketplace add thedotmack/claude-mem and /plugin install claude-mem inside Claude Code, then restart.

Who performs the memory compression, and what does it cost?

The README lists the claude-mem observer (14-day free trial, then fallback to your Anthropic plan unless you subscribe), your own OpenRouter or Gemini key, or your Anthropic plan. Compression calls a language model, so the cost comes from plan quota or third-party billing.

How does it capture data on hosts without hooks?

The README names Grok Bot as the example: with no host hooks, it watches the chat log files. This needs no host event API, but it reads after the fact, so freshness depends on the log format.