On the edge of Apple Silicon memory speeds
Details
- External ID
- 46656142
- Source
- HN
- Company
- —
- Product
- On the edge of Apple Silicon memory speeds
- Website domain
- github.com
- Launched
- Jan. 17, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.09617918313570488
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
I have developed open source CLI-tool for Apple Silicon macOS. It measures memory speeds in different ways and also latency. It can achieve up to 96-97% efficiency on read speed on M4 base what is advertised as 120GB/s. All memory operations are in assembly.I would really appreciate for results on different CPU's how benchmark works on those. I have been able to test this on M1 and M4.command : 'memory_benchmark -non-cacheable -count 5 -output results.JSON' (close all applications before running)This will generate JSON file where you find sections copy_gb_s, read_gb_s and write_gb_s statics.Example M4 with 10 loops: "copy_gb_s": { "statistics": { "average": 106.65421233311835, "max": 106.70240696071005, "median": 106.65069297260811, "min": 106.6336774994254, "p90": 106.66606919223108, "p95": 106.68423807647056, "p99": 106.69877318386216, "stddev": 0.01930653530818627 }, "values": [ 106.70240696071005, 106.66203166240008, 106.64410802226159, 106.65831409449595, 106.64148106986977, 106.6482935780762, 106.63974821679058, 106.65896986001393, 106.6336774994254, 106.65309236714002 ] }, "read_gb_s": { "statistics": { "average": 115.83111228356601, "max": 116.11098114619033, "median": 115.84480882265643, "min": 115.56959026587722, "p90": 115.99667266786554, "p95": 116.05382690702793, "p99": 116.09955029835784, "stddev": 0.1768243167963439 }, "values": [ 115.79154681380165, 115.56959026587722, 115.60574235736468, 115.72112860271632, 115.72147129262802, 115.89807083151123, 115.95527337086908, 115.95334642887214, 115.98397172582945, 116.11098114619033 ] }, "write_gb_s": { "statistics": { "average": 65.55966046805113, "max": 65.59040040480241, "median": 65.55933583741347, "min": 65.50911885624045, "p90": 65.5840272860955, "p95": 65.58721384544896, "p99": 65.58976309293172, "stddev": 0.02388146120866979 },Patterns benchmark also shows bit more of memory speeds. command: 'memory_benchmark -patterns -non-cacheable -count 5 -output patterns.JSON'Example M4 from 100 loops: "sequential_forward": { "bandwidth": { "read_gb_s": { "statistics": { "average": 116.38363691482549, "max": 116.61212708384109, "median": 116.41264548721367, "min": 115.449510036971, "p90": 116.54143114134801, "p95": 116.57314206456576, "p99": 116.60095068065866, "stddev": 0.17026641589059727 } } } }"strided_4096": { "bandwidth": { "read_gb_s": { "statistics": { "average": 26.460392735220456, "max": 27.7722419653915, "median": 26.457051473208285, "min": 25.519925729459107, "p90": 27.105171215736604, "p95": 27.190715938337473, "p99": 27.360449534513144, "stddev": 0.4730857335572576 } } } }"random": { "bandwidth": { "read_gb_s": { "statistics": { "average": 26.71367836895143, "max": 26.966820487564327, "median": 26.69907406197067, "min": 26.49374804466308, "p90": 26.845236287807374, "p95": 26.882004355057887, "p99": 26.95742242818151, "stddev": 0.09600564296001704 } } } }Thank you for reading :)
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- apple silicon memory performance exploration
- Manually corrected
- False
Could you build this?
No Benchmarking low-level memory latency and bandwidth near theoretical hardware limits requires hand-written ARM64 assembly, cache-line eviction tricks, and deep knowledge of Apple Silicon microarchitecture.
What it would actually take: Building this requires hand-crafted ARM64 assembly routines utilizing NEON/AMX vector instructions, unrolled load/store loops, memory barriers (`dmb`, `isb`), and hardware performance counters via macOS kernel APIs or `mach_time`. The hard part is avoiding CPU pipeline stalls, tuning cache prefetchers, managing virtual-to-physical address mappings, and understanding the specific memory controller topologies of M-series chips. This demands specialized microarchitecture and low-level systems programming expertise.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 149 launches clear the similarity bar, closest 8 shown.
Attention rank: #146 of 150 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 74 days after the earliest competitor.
- I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac · hn · 2026-08-16 · 21 upvotes · similarity 0.52
- Shoehorn, a library to quantize an LLM to fit your Mac's VRAM · hn · 2026-08-14 · 6 upvotes · similarity 0.44
- Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU · hn · 2026-06-30 · 10 upvotes · similarity 0.42
- Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac · hn · 2026-07-29 · 919 upvotes · similarity 0.42
- The Type 1 Civilization Toolkit · ph · 2026-09-24 · 1 upvotes · similarity 0.42
- Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone · hn · 2026-08-03 · 312 upvotes · similarity 0.42
- m6502, a 6502 CPU for FPGAs and Tiny Tapeout · hn · 2026-02-18 · 5 upvotes · similarity 0.41
- Mstat a temperature and stat tracker for macOS in Zig and SwiftUI · hn · 2026-08-03 · 13 upvotes · similarity 0.41
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.