ClawMem
Open-source agent memory with SOTA local GPU retrieval
Details
- External ID
- 47472965
- Source
- HN
- Company
- —
- Product
- ClawMem
- Website domain
- github.com
- Launched
- March 22, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.1070110701107011
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
So I've been building ClawMem, an open-source context engine that gives AI coding agents persistent memory across sessions. It works with Claude Code (hooks + MCP) and OpenClaw (ContextEngine plugin + REST API), and both can share the same SQLite vault, so your CLI agent and your voice/chat agent build on the same memory without syncing anything.The retrieval architecture is a Frankenstein, which is pretty much always my process. I pulled the best parts from recent projects and research and stitched them together: [QMD](https://github.com/tobi/qmd) for the multi-signal retrieval pipeline (BM25 + vector + RRF + query expansion + cross-encoder reranking), [SAME](https://github.com/sgx-labs/statelessagent) for composite scoring with content-type half-lives and co-activation reinforcement, [MAGMA](https://arxiv.org/abs/2501.13956) for intent classification with multi-graph traversal (semantic, temporal, and causal beam search), [A-MEM](https://arxiv.org/abs/2510.02178) for self-evolving memory notes, and [Engram](https://github.com/Gentleman-Programming/engram) for deduplication patterns and temporal navigation. None of these were designed to work together. Making them coherent was most of the work.On the inference side, QMD's original stack uses a 300MB embedding model, a 1.1GB query expansion LLM, and a 600MB reranker. These run via llama-server on a GPU or in-process through node-llama-cpp (Metal, Vulkan, or CPU). But the more interesting path is the SOTA upgrade: ZeroEntropy's distillation-paired zembed-1 + zerank-2. These are currently the top-ranked embedding and reranking models on MTEB, and they're designed to work together. The reranker was distilled from the same teacher as the embedder, so they share a semantic space. You need ~12GB VRAM to run both, but retrieval quality is noticeably better than the default stack. There's also a cloud embedding option if you're tight on vram or prefer to offload embedding to a cloud model.For Claude Code specifically, it hooks into lifecycle events. Context-surfacing fires on every prompt to inject relevant memory, decision-extractor and handoff-generator capture session state, and a feedback loop reinforces notes that actually get referenced. That handles about 90% of retrieval automatically. The other 10% is 28 MCP tools for explicit queries. For OpenClaw, it registers as a ContextEngine plugin with the same hook-to-lifecycle mapping, plus 5 REST API tools for the agent to call directly.It runs on Bun with a single SQLite vault (WAL mode, FTS5 + vec0). Everything is on-device; no cloud dependency unless you opt into cloud embedding. The whole system is self-contained.This is a polished WIP, not a finished product. I'm a solo dev. The codebase is around 19K lines and the main store module is a 4K-line god object that probably needs splitting. And of course, the system is only as good as what you index. A vault with three memory files gives deservedly thin results. One with your project docs, research notes, and decision records gives something actually useful.Two questions I'd genuinely like input on: (1) Has anyone else tried running SOTA embedding + reranking models locally for agent memory, and is the quality difference worth the VRAM? (2) For those running multiple agent interfaces (CLI + voice/chat), how are you handling shared memory today?
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Hobby / open-source project
- Normalized one-liner
- agent memory system with gpu retrieval
- Manually corrected
- False
Could you build this?
Partial Integrating an MCP server with SQLite is simple, but achieving true SOTA local GPU vector retrieval and effective semantic memory architectures requires specialized ML engineering.
What it would actually take: A daemon written in Rust or Python utilizing ONNX Runtime or LibTorch with CUDA acceleration for local embeddings and cross-encoder re-ranking. The hard part is managing GPU memory footprints, low-latency concurrent retrieval under local resource constraints, and designing a robust agent memory compaction and retrieval policy.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 225 launches clear the similarity bar, closest 8 shown.
Attention rank: #205 of 226 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 143 days after the earliest competitor.
- PMB · hn · 2026-06-22 · 7 upvotes · similarity 0.57
- Remembrane · hn · 2026-08-07 · 13 upvotes · similarity 0.56
- A file-based agent memory framework that works like skill · hn · 2026-01-06 · 11 upvotes · similarity 0.51
- Memori · ph · 2026-05-28 · 168 upvotes · similarity 0.51
- Mwe-MCP · hn · 2026-07-23 · 5 upvotes · similarity 0.51
- Unify memory across agents and improve context rot, written in Rust · hn · 2026-04-05 · 5 upvotes · similarity 0.51
- CoreMem · hn · 2026-05-22 · 5 upvotes · similarity 0.50
- Slowave · hn · 2026-09-14 · 5 upvotes · similarity 0.49
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.