Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Running PrismML's Bonsai inside DRAM by breaking DDR4 timing rules

Details

External ID
49019271
Source
HN
Company
—
Product
—
Website domain
—
Launched
July 23, 2026
Cohort
—
Upvotes
23
Upvotes percentile
0.7437275985663082
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

The excitement surrounding PrismML’s 1-bit/ternary Bonsai models has the industry closely watching how smartphone giants, particularly Apple, will implement LLMs on edge devices.Moving AI on-device is a brilliant and necessary strategy. It ensures absolute user privacy in alignment with EU regulations, fundamentally shifts the economics away from costly cloud inference, and paves the way for a significant hardware upgrade supercycle as users seek true AI-capable silicon.To create a smart on-device "Semantic Router," models need to reach the 27B+ parameter scale. Achieving this on a phone requires extreme quantization, such as PrismML’s ternary weights.However, a critical hardware reality often overlooked by the software world is that fitting the weights in RAM is not equivalent to moving them. Running a 27B ternary model on standard LPDDR encounters a significant memory bandwidth limitation. Transferring gigabytes of data across the SoC bus for each token generation can lead to thermal throttling of the NPU and excessive battery drain.This raises an important question: why are we still transferring data to the compute? Why not execute AI inference natively within the memory?Frustrated with academic PIM simulations that overlook bare-metal physics, I developed CaSA, an architecture that performs ternary LLM inference directly inside COTS DRAM through charge-sharing, completely bypassing the memory bus.Software quantization is a great initial step, and CaSA provides the physical hardware substrate needed to complete the bridge: https://github.com/pcdeni/CaSA

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI feature
Project type
Hobby / open-source project
Normalized one-liner
optimize ml model memory using dram techniques
Manually corrected
False

Could you build this?

No Operating neural network inference directly in DRAM by breaking DDR4 timing parameters involves low-level hardware memory controller exploitation, electrical engineering, and physical silicon properties that software vibe coding cannot touch.

What it would actually take: This requires low-level C, assembly, and custom kernel/FPGA drivers capable of bypassing standard JEDEC DDR4 memory timing rules (row-hammer style or custom memory controller microcode) to perform in-memory processing. Developers need deep expertise in memory bus electrical signaling, JEDEC hardware specifications, and quantization-aware machine learning compiler design.

Discussion

10 comments analyzed.

Competitors mentioned: GPU-based inference, Traditional off-the-shelf memory systems

Concerns raised: Not faster than GPU for practical use, 47.5 seconds per token performance is slow, Requires silicon-level changes beyond prototype stage, Claims of 13x speedup lack proof, Poor documentation and AI-generated prose quality

Feature requests: Dynamic voltage adjustment per memory cell groups, Human-written documentation and descriptions, Clear setup instructions and benchmarking details

Competitors

Other products that read as similar to this one — 145 launches clear the similarity bar, closest 8 shown.

Attention rank: #33 of 146 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 257 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.