Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

L88

A Local RAG System on 8GB VRAM (Need Architecture Feedback)

Details

External ID
47133027
Source
HN
Company
—
Product
—
Website domain
—
Launched
Feb. 24, 2026
Cohort
—
Upvotes
12
Upvotes percentile
0.5970350404312669
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hey everyone,I’ve been working on a project called L88 — a local RAG system that I initially focused on UI/UX for, so the retrieval and model architecture still need proper refinement.Repo: https://github.com/Hundred-Trillion/L88-FullI’m running this on 8GB VRAM and a strong CPU (128GB RAM). Embeddings and preprocessing run on CPU, and the main model runs on GPU. One limitation I ran into is that my evaluator and generator LLM ended up being the same model due to compute constraints, which defeats the purpose of evaluation.I’d really appreciate feedback on:Better architecture ideas for small-VRAM RAGSplitting evaluator/generator roles effectivelyImproving the LangGraph pipelineAny bugs or design smells you noticeWays to optimize the system for local hardwareI’m 18 and still learning a lot about proper LLM architecture, so any technical critique or suggestions would help me grow as a developer. If you check out the repo or leave feedback, it would mean a lot — I’m trying to build a solid foundation and reputation through real projects.Thanks!

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
local rag system for low-vram systems
Manually corrected
False

Could you build this?

Yes This is a local RAG interface built on standard open-source tools (like Ollama, LangChain/LlamaIndex, and local vector stores) constrained to run on 8GB VRAM.

Discussion

3 comments analyzed.

Competitors mentioned: Qdrant, FAISS, SQLite, Supabase (vector/storage solutions), BGE-base-en-v1.5, bge-reranker-v2-m3 (embedding/reranking models), Cross-encoder rerankers, BM25 (keyword search)

Concerns raised: VRAM/RAM constraints on free-tier/low-resource servers, Latency from multiple sequential LLM calls, Circular evaluation issue: using same model for retrieval and quality scoring, Vector search alone misses exact-match queries (titles, authors), Cross-encoder rerankers destroy performance on resource-constrained setups

Feature requests: Merge router/analyzer/rewriter into single LLM call to reduce latency, Implement hybrid BM25 + vector search with dynamic weighting by query type, Add semantic caching layer (cosine similarity threshold) alongside exact-match hashing, Use cross-encoder for evaluation/thresholding instead of LLM-as-judge, Parent-child chunking: embed small chunks, store large payloads in lightweight DB

Competitors

Other products that read as similar to this one — 172 launches clear the similarity bar, closest 8 shown.

Attention rank: #82 of 173 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 116 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.