Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Unified multimodal memory framework, without embeddings

Details

External ID
46525174
Source
HN
Company
—
Product
Unified multimodal memory framework, without embeddings
Website domain
github.com
Launched
Jan. 7, 2026
Cohort
—
Upvotes
7
Upvotes percentile
0.3544137022397892
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hi HN,We’ve been building memU(https://github.com/NevaMind-AI/memU), an open-source, general-purpose memory framework for AI agents. It supports dual-mode retrieval: classic RAG and LLM-based direct file reading.Most multimodal memory systems either embed everything into vectors or treat non-text data as attachments. These work, but at scale it becomes hard to explain why certain context was retrieved and what evidence it relies on.memU takes a different approach: since models reason in language, multimodal memory should converge into structured, queryable text, while remaining fully traceable to original data.---## Three-Layer Architecture- Resource Layer Stores raw multimodal data as ground truth. All higher-level memory remains traceable to this layer.- Memory Item Layer Extracts atomic facts from raw data and stores them as natural-language statements. Embeddings are optional and used only for acceleration.- Memory Category Layer Aggregates items into readable, theme-based memory files (e.g. user preferences, work logs). Frequently accessed topics stay active; low-usage content is demoted to balance speed and coverage.---## Memorization Bottom-up and asynchronous. Data flows from resources → items → category files without manual schemas. When capacity is reached, recently relevant memories replace the least used ones.## Retrieval Top-down. memU searches category files first, then items, and only falls back to raw data if needed. At the item layer, it combines BM25 + embeddings to balance exact matching and semantic recall, avoiding embedding-only imprecision.Dual-mode retrieval lets applications choose between: - low-latency embedding search, or - LLM-based direct reading of memory files.## Evolution Memory structure adapts automatically based on real usage: - Frequently accessed memories remain at the Category layer - Memories retrieved from raw data are promoted upward and linked - Organization evolves from usage patterns, not predefined rulesGoal: keep relevant memories retrievable at the Category layer and minimize latency over time.---## A Unified Multimodal Memory Pipeline memU is a text-centered multimodal memory system. Multimodal inputs are progressively converted into interpretable text memory, while staying traceable to original data. This provides stable, high-level context for reasoning, with detailed evidence available when needed—inside a memory structure that evolves through real-world use.

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
multimodal memory framework without embeddings
Manually corrected
False

Could you build this?

Partial Basic multimodal orchestration can be built with LLMs, but creating an efficient multimodal memory framework without embeddings requires designing novel document indexing and fast context-filtering algorithms.

What it would actually take: The system requires a hybrid architecture integrating fast OCR/vision pipelines, structured hierarchical file-system indexing, and custom prompt-driven or rule-based semantic filtering to avoid embedding costs without blowing context budgets. Hard parts include managing latency when scanning raw multimodal files and ensuring high recall across diverse formats (images, audio, code) without traditional vector search.

Discussion

2 comments analyzed.

Competitors mentioned: Vanilla vector search / standard RAG, Embedding-based memory stacks

Concerns raised: Retrieval accuracy improvement over standard RAG not quantified, Unclear performance 'lift' for migration decisions, Scalability of human-readable memory format for long-running agents

Feature requests: Benchmarks comparing retrieval accuracy to vanilla vector search, Temporal reasoning and traceability capabilities, Audit/inspection tooling for agent memory state

Competitors

Other products that read as similar to this one — 114 launches clear the similarity bar, closest 8 shown.

Attention rank: #65 of 115 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 68 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.