Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

RunMat

runtime with auto CPU/GPU routing for dense math

Details

External ID
46121951
Source
HN
Company
—
Product
RunMat
Website domain
github.com
Launched
Dec. 2, 2025
Cohort
—
Upvotes
21
Upvotes percentile
0.6612595419847328
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hi, I’m Nabeel. In August I released RunMat as an open-source runtime for MATLAB code that was already much faster than GNU Octave on the workloads I tried. https://news.ycombinator.com/item?id=44972919Since then, I’ve taken it further with RunMat Accelerate: the runtime now automatically fuses operations and routes work between CPU and GPU. You write MATLAB-style code, and RunMat runs your computation across CPUs and GPUs for speed. No CUDA, no kernel code.Under the hood, it builds a graph of your array math, fuses long chains into a few kernels, keeps data on the GPU when that helps, and falls back to CPU JIT / BLAS for small cases.On an Apple M2 Max (32 GB), here are some current benchmarks (median of several runs):* 5M-path Monte Carlo * RunMat ≈ 0.61 s * PyTorch ≈ 1.70 s * NumPy ≈ 79.9 s → ~2.8× faster than PyTorch and ~130× faster than NumPy on this test.* 64 × 4K image preprocessing pipeline (mean/std, normalize, gain/bias, gamma, MSE) * RunMat ≈ 0.68 s * PyTorch ≈ 1.20 s * NumPy ≈ 7.0 s → ~1.8× faster than PyTorch and ~10× faster than NumPy.* 1B-point elementwise chain (sin / exp / cos / tanh mix) * RunMat ≈ 0.14 s * PyTorch ≈ 20.8 s * NumPy ≈ 11.9 s → ~140× faster than PyTorch and ~80× faster than NumPy.If you want more detail on how the fusion and CPU/GPU routing work, I wrote up a longer post here: https://runmat.org/blog/runmat-accel-intro-blogYou can run the same benchmarks yourself from the GitHub repo in the main HN link. Feedback, bug reports, and “here’s where it breaks or is slow” examples are very welcome.

Enrichment

Theme
gpu compute and acceleration tools
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
Not AI
Project type
Commercial product
Normalized one-liner
runtime with automatic cpu/gpu routing for dense math
Manually corrected
False

Could you build this?

No Developing a custom high-performance runtime for MATLAB that automatically performs operator fusion and dynamic CPU/GPU routing requires deep expertise in compiler design, numerical computing, and GPU execution models.

What it would actually take: A production implementation requires a language frontend (AST parser/type inferencer for MATLAB semantics), an intermediate representation (like MLIR), and an optimizing compiler pipeline that handles JIT compilation, operator fusion, and memory management across host and device memory (CUDA/ROCm/Metal). It demands specialized knowledge in high-performance linear algebra (BLAS/LAPACK), hardware memory architectures, and compiler construction.

Discussion

4 comments analyzed.

Competitors mentioned: NumPy, PyTorch, Octave, Julia, MATLAB

Concerns raised: Unclear target audience (MATLAB vs Python vs Julia users), Whether it actually outperforms custom CUDA/GPU code, Market size for MATLAB-compatible runtime

Competitors

Other products that read as similar to this one — 153 launches clear the similarity bar, closest 8 shown.

Attention rank: #66 of 154 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 34 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.