Sparse Matrix-Vector Multiplication that works at 30–90% sparsity
Details
- External ID
- 46046106
- Source
- HN
- Company
- —
- Product
- Sparse Matrix-Vector Multiplication that works at 30–90% sparsity
- Website domain
- github.com
- Launched
- Nov. 25, 2025
- Cohort
- —
- Upvotes
- 7
- Upvotes percentile
- 0.37882096069869
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
To get benefits from sparsity, you usually need to have very sparse matrices, impose some structure on the sparsity pattern or have specialized hardware. None of it is the case if you want to rune pruned LLMs on consumer devices. I wanted to see how far can you push it on a GPU and ended up with this. Blog: https://www.grizzlytech.dev/blog/macko-spmv Paper: https://arxiv.org/abs/2511.13061 Code (example with torch): https://github.com/vlejd/macko_spmv
Enrichment
- Theme
- gaming performance and optimization utilities
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- sparse matrix-vector multiplication optimization
- Manually corrected
- False
Could you build this?
No Engineering sparse matrix-vector multiplication that outperforms dense kernels at intermediate 30-90% sparsity requires deep mathematical formulation and manual GPU hardware microarchitecture optimization.
What it would actually take: The architecture relies on custom CUDA or Triton GPU kernels targeting modern tensor or compute cores with specialized compression formats (such as blocked CSR, sliced ELLPACK, or custom bitmasks). The hard problem is eliminating warp divergence and coalescing irregular memory access patterns when sparsity is too low for traditional sparse BLAS to win over dense GEMM. Requires advanced GPU performance engineering, assembly/PTX-level optimization, and deep numerical linear algebra expertise.
Discussion
5 comments analyzed.
Competitors mentioned: cuBLAS, Quantization (fp16 to fp8)
Concerns raised: Unstructured sparsity lacks efficient GPU kernel support, Structured sparsity degrades model quality, Quantization provides easier speedup than sparsity, Limited applicability outside 30-90% sparsity range, Block sparsity creates worst-case scenarios
Feature requests: Support for use cases beyond 30-90% sparsity, Guidance on wider adoption of sparse neural approaches
Competitors
Other products that read as similar to this one — 37 launches clear the similarity bar, closest 8 shown.
Attention rank: #29 of 38 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 24 days after the earliest competitor.
- RunMat · hn · 2025-12-02 · 21 upvotes · similarity 0.37
- Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO) · hn · 2026-08-01 · 21 upvotes · similarity 0.36
- Axiom · hn · 2026-02-02 · 5 upvotes · similarity 0.36
- The Hessian of tall-skinny networks is easy to invert · hn · 2026-01-15 · 31 upvotes · similarity 0.36
- EdgeVec · hn · 2025-12-12 · 7 upvotes · similarity 0.35
- Run TRELLIS.2 Image-to-3D generation natively on Apple Silicon · hn · 2026-04-20 · 202 upvotes · similarity 0.35
- FastVSR · ph · 2026-09-15 · 1 upvotes · similarity 0.35
- TIRx-kernels · github · 2026-09-29 · 16 upvotes · similarity 0.34
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.