NanoEuler
GPT-2 scale model in pure C/CUDA from scratch
Details
- External ID
- 48710778
- Source
- HN
- Company
- —
- Product
- NanoEuler
- Website domain
- github.com
- Launched
- June 28, 2026
- Cohort
- —
- Upvotes
- 55
- Upvotes percentile
- 0.8586065573770492
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hi everyone,I started working on nanoeuler after the ban of anthropic's fable because my ambition and dream is to work in the AI field in anthropic. The two interesting reasons that led me to create nanoeuler were (1) interfacing with llm does not mean understanding how they are composed and (2), working on llm with a very low-level layer to understand the correlation between parameters and data and growth of the model and how the GPU works and how some layers can be optimized.So I started working on it with a research aspect by making nanoeuler grow more and more but doing one step after another starting from Shakespeare.txt and understanding what a text generation model understands at 23 million parameters. For example, nanoeuler at that number had understood that Name: started a line and wrote that line with sense.I wrote everything in CUDA because I wanted to not use any intermediary between the model in training and inference and what it had to do. Then the use of SFT and much more, even if in small ways, were really useful to understand the various step to make an llm like a chatbot.Any feedback, help, or suggestions are absolutely welcome!
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- gpt-2 scale model implementation in pure c/cuda
- Manually corrected
- False
Could you build this?
No Implementing a GPT-2 architecture from scratch in pure C and CUDA requires deep knowledge of low-level parallel computing, GPU memory management, and neural network mathematical fundamentals.
What it would actually take: This project requires writing custom CUDA kernels for matrix multiplications, attention mechanisms, layer normalization, and backward-pass gradients without high-level abstractions like PyTorch. The developer must manually optimize shared memory usage, tensor core utilization, warp synchronization, and memory bandwidth, requiring specialized GPU systems engineering skills.
Discussion
20 comments analyzed.
Competitors mentioned: PyTorch, cuBLAS
Concerns raised: Performance 1.5-2.5x slower than PyTorch, Kernel launch overhead and memory traffic inefficiency, README appears AI-generated, Sparse commit history on projects, Unclear how much code is LLM-generated
Feature requests: Author disclosure requirement for LLM usage in projects, Further kernel optimization and aggressive kernel merging, Improved documentation clarity and readability
Competitors
Other products that read as similar to this one — 88 launches clear the similarity bar, closest 8 shown.
Attention rank: #18 of 89 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 240 days after the earliest competitor.
- GPT‑5.4 mini and nano · ph · 2026-03-18 · 267 upvotes · similarity 0.48
- Revibing nanochat's inference model in C++ with ggml · hn · 2026-01-10 · 5 upvotes · similarity 0.46
- WebGCM · hn · 2026-09-21 · 6 upvotes · similarity 0.44
- Interactive 3D and 2D Visualization of GPT-2 · hn · 2026-03-19 · 12 upvotes · similarity 0.43
- Mini-AGI · hn · 2026-09-21 · 277 upvotes · similarity 0.43
- I built GPT from scratch to understand how it works · hn · 2026-01-14 · 7 upvotes · similarity 0.43
- Tiny Diffusion · hn · 2025-11-10 · 172 upvotes · similarity 0.41
- MicroGPT in 243 Lines · hn · 2026-02-13 · 10 upvotes · similarity 0.40
Other launches for this product
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.