L88
A Local RAG System on 8GB VRAM (Need Architecture Feedback)
Details
- External ID
- 47133027
- Source
- HN
- Company
- —
- Product
- —
- Website domain
- —
- Launched
- Feb. 24, 2026
- Cohort
- —
- Upvotes
- 12
- Upvotes percentile
- 0.5970350404312669
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hey everyone,I’ve been working on a project called L88 — a local RAG system that I initially focused on UI/UX for, so the retrieval and model architecture still need proper refinement.Repo: https://github.com/Hundred-Trillion/L88-FullI’m running this on 8GB VRAM and a strong CPU (128GB RAM). Embeddings and preprocessing run on CPU, and the main model runs on GPU. One limitation I ran into is that my evaluator and generator LLM ended up being the same model due to compute constraints, which defeats the purpose of evaluation.I’d really appreciate feedback on:Better architecture ideas for small-VRAM RAGSplitting evaluator/generator roles effectivelyImproving the LangGraph pipelineAny bugs or design smells you noticeWays to optimize the system for local hardwareI’m 18 and still learning a lot about proper LLM architecture, so any technical critique or suggestions would help me grow as a developer. If you check out the repo or leave feedback, it would mean a lot — I’m trying to build a solid foundation and reputation through real projects.Thanks!
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- local rag system for low-vram systems
- Manually corrected
- False
Could you build this?
Yes This is a local RAG interface built on standard open-source tools (like Ollama, LangChain/LlamaIndex, and local vector stores) constrained to run on 8GB VRAM.
Discussion
3 comments analyzed.
Competitors mentioned: Qdrant, FAISS, SQLite, Supabase (vector/storage solutions), BGE-base-en-v1.5, bge-reranker-v2-m3 (embedding/reranking models), Cross-encoder rerankers, BM25 (keyword search)
Concerns raised: VRAM/RAM constraints on free-tier/low-resource servers, Latency from multiple sequential LLM calls, Circular evaluation issue: using same model for retrieval and quality scoring, Vector search alone misses exact-match queries (titles, authors), Cross-encoder rerankers destroy performance on resource-constrained setups
Feature requests: Merge router/analyzer/rewriter into single LLM call to reduce latency, Implement hybrid BM25 + vector search with dynamic weighting by query type, Add semantic caching layer (cosine similarity threshold) alongside exact-match hashing, Use cross-encoder for evaluation/thresholding instead of LLM-as-judge, Parent-child chunking: embed small chunks, store large payloads in lightweight DB
Competitors
Other products that read as similar to this one — 172 launches clear the similarity bar, closest 8 shown.
Attention rank: #82 of 173 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 116 days after the earliest competitor.
- I wrote a custom assembler for CHIP-8 in C++ · hn · 2026-09-17 · 28 upvotes · similarity 0.49
- E80: an 8-bit CPU in structural VHDL · hn · 2026-01-17 · 34 upvotes · similarity 0.49
- Shoehorn, a library to quantize an LLM to fit your Mac's VRAM · hn · 2026-08-14 · 6 upvotes · similarity 0.49
- VRAMGlass · ph · 2026-09-19 · 1 upvotes · similarity 0.47
- Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO) · hn · 2026-08-01 · 21 upvotes · similarity 0.46
- I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp · hn · 2026-07-29 · 5 upvotes · similarity 0.44
- LocalLLM · hn · 2026-04-23 · 16 upvotes · similarity 0.43
- AMP · hn · 2025-12-14 · 5 upvotes · similarity 0.43
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.