Running BitNet b1.58 inside DRAM by breaking DDR4 timing rules
Details
- External ID
- 48250231
- Source
- HN
- Company
- —
- Product
- —
- Website domain
- —
- Launched
- May 23, 2026
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.3053311793214863
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I have been working on running BitNet b1.58 inside DRAM by intentionally breaking DDR4 timing rules. Also made a visual explainer: https://pcdeni.github.io/CaSA/explainer/ This is tested and works inside commercial off the shelf memory with custom memory controller in the FPGA. The underlying effect is well characterized in academic papers (cmu safari, simra, dram bender, etc). In the process of getting this to work I also made previously undocumented discovery about DDR behaviour: https://pcdeni.github.io/CaSA/explainer/xor-spread.html Overall it is a bit slow, since data (in full rows) needs to be moved even when what is actually needed is only the count of the '1' bits (popcount). To make it competitive memory die changes would be needed, but not as drastic as merging compute and memory into one silicon. This would then avoid the memory wall issue the industry is currently facing.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- efficient llm inference in dram
- Manually corrected
- False
Could you build this?
No Running quantized neural networks directly inside DRAM chips via DDR4 timing rule violations requires specialized electrical engineering, custom FPGA memory controller RTL, and hardware-level validation.
What it would actually take: Building this requires custom Verilog/VHDL RTL memory controllers running on an FPGA wired to DDR4 DIMM slots to issue out-of-spec command sequences at sub-nanosecond precision. The hard parts are exploiting analog charge sharing across DRAM bitlines for in-situ computation and handling high hardware bit-error rates. This demands specialized domain knowledge in DRAM internal architectures, electrical signal integrity, and hardware verification equipment.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 44 launches clear the similarity bar, closest 8 shown.
Attention rank: #30 of 45 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 192 days after the earliest competitor.
- 8Gb DDR4 DRAM Chip · Wide-temp · CN · ph · 2026-09-17 · 1 upvotes · similarity 0.51
- We built an 8-bit CPU as 2nd year EE students · hn · 2026-06-15 · 108 upvotes · similarity 0.42
- Industrial DDR3 DRAM Chip · ph · 2026-09-16 · 1 upvotes · similarity 0.41
- E80: an 8-bit CPU in structural VHDL · hn · 2026-01-17 · 34 upvotes · similarity 0.40
- Running PrismML's Bonsai inside DRAM by breaking DDR4 timing rules · hn · 2026-07-23 · 23 upvotes · similarity 0.40
- CodeYam Memory · hn · 2026-03-04 · 17 upvotes · similarity 0.37
- CoreTrace, a visual 16-bit CPU simulator · hn · 2026-08-13 · 5 upvotes · similarity 0.36
- A nibble-oriented CPU in Verilog to build a scientific calculator · hn · 2026-05-15 · 119 upvotes · similarity 0.35
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.