Autograd.c
A tiny ML framework built from scratch
Details
- External ID
- 46285424
- Source
- HN
- Company
- —
- Product
- Autograd.c
- Website domain
- github.com
- Launched
- Dec. 16, 2025
- Cohort
- —
- Upvotes
- 85
- Upvotes percentile
- 0.8807251908396947
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
built a tiny pytorch clone in c after going through prof. vijay janapa reddi's mlsys book: mlsysbook.ai/tinytorch/perfect for learning how ml frameworks work under the hood :)
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- minimal machine learning framework
- Manually corrected
- False
Could you build this?
No Writing a tensor autograd framework from scratch in C demands deep mathematical mastery of reverse-mode automatic differentiation, tensor stride manipulation, and low-level memory management.
What it would actually take: A pure C implementation requires custom memory allocators for dynamically sized n-dimensional tensors, a computational directed acyclic graph (DAG) engine supporting tape-based backpropagation, and optimized BLAS-like numerical kernels. Deep expertise in computer systems, linear algebra, memory alignment, and compiler-friendly SIMD vectorization is required.
Discussion
13 comments analyzed.
Competitors mentioned: Enzyme (LLVM IR level autograd), JAX (architecture/design reference), Theano (compiler-autograd method), MLIR/IREE (heap-free implementation direction)
Concerns raised: Compilation cost O(n) times per forward/reverse pass, Finding ideal forward vs reverse mode combination is NP-hard, Stack pressure and memory issues with higher-order operators, Forward mode alone doubles stack pressure; reverse mode needs 2x stack, Diminishes C's performance advantages with indirection overhead
Feature requests: In-place gradient accumulation instead of creating new tensors, Heap-free implementation for memory efficiency, Recompile per iteration when weights updated at batch/distributed intervals
Competitors
Other products that read as similar to this one — 52 launches clear the similarity bar, closest 8 shown.
Attention rank: #8 of 53 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 43 days after the earliest competitor.
- TrenTorch · ph · 2026-09-20 · 2 upvotes · similarity 0.48
- Reimplementing PyTorch from scratch (MLP, CNN) to learn the internals · hn · 2026-02-03 · 5 upvotes · similarity 0.45
- Flow Matching model inference in C · hn · 2026-07-23 · 5 upvotes · similarity 0.43
- Microcrad · hn · 2026-06-17 · 79 upvotes · similarity 0.42
- CJIT, a single-binary C compiler that can self host · hn · 2026-04-13 · 7 upvotes · similarity 0.42
- Minimal DL library in C · hn · 2025-12-17 · 13 upvotes · similarity 0.38
- Linear RNN/Reservoir hybrid generative model, one C file (no deps.) · hn · 2026-04-09 · 7 upvotes · similarity 0.38
- Self-growing neural networks via a custom Rust-to-LLVM compiler · hn · 2025-12-28 · 8 upvotes · similarity 0.37
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.