Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Llm.sql

Run a 640MB LLM on SQLite, with 210MB peak RSS and 7.4 tok/s

Details

External ID
47888712
Source
HN
Company
—
Product
—
Website domain
—
Launched
April 24, 2026
Cohort
—
Upvotes
8
Upvotes percentile
0.4813624678663239
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hi HN,I built llm.sql, an LLM inference framework that reimagines the LLM execution pipeline as a series of structured SQL queries atop SQLite.The motivation: Edge LLMs are getting better, but hardware remains a bottleneck, especially RAM (size and bandwidth).When available memory is less than the model size and KV cache, the OS incurs page faults and swaps pages using LRU-like strategies, resulting in throughput degradation that's hard to notice and even harder to debug. In fact, the memory access pattern during LLM inference is deterministic - we know exactly which weights are needed and when. This means even Bélády's optimal page replacement algorithm is applicable here.So instead of letting the OS manage memory, llm.sql takes over:- Model parameters are stored in SQLite BLOB tables- Computational logic is implemented as SQLite C extensions- Memory management is handled explicitly, not by the OS- Zero heavy dependencies. No PyTorch, no Transformers. Just Python, C, or C++This gives us explicit, deterministic control over what's in memory at each step of inference.Results:Running Qwen2.5-0.5B-INT8 (~640MB model) with a peak RSS ~210MB and 7.40 tokens/s throughput.Alpha version is available on GitHub: https://github.com/xuxianghong12/llm.sqlI'm the developer, happy to answer any technical questions about the design and implementation.

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
lightweight llm on sqlite
Manually corrected
False

Could you build this?

No Writing an LLM inference engine where transformer forward-pass tensor operations and matrix multiplications are reimagined as native SQL queries inside SQLite requires deep systems and deep learning runtime engineering.

What it would actually take: The core engine requires custom SQLite C extensions or virtual tables with deeply optimized BLAS/vectorized kernels (SIMD, AVX/NEON) to perform matrix math, KV-caching, and layer-by-layer attention inside SQLite query plans. The hard parts are memory-mapping model weights in low RAM footprints (210MB RSS), avoiding intermediate buffer allocations across relational operations, and achieving non-trivial token speeds (7+ tok/s). This demands elite expertise in database internals, systems programming (C/Rust), and low-level ML compiler/runtime architecture.

Discussion

2 comments analyzed.

Concerns raised: Project still in alpha, validation phase ongoing, Uncertain if suitable for Raspberry Pi despite theoretical possibility

Feature requests: Benchmark across different hardware platforms, Test and optimize for edge devices like Raspberry Pi

Competitors

Other products that read as similar to this one — 243 launches clear the similarity bar, closest 8 shown.

Attention rank: #132 of 244 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 177 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.