Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

NanoEuler

GPT-2 scale model in pure C/CUDA from scratch

Details

External ID
48601472
Source
HN
Company
—
Product
NanoEuler
Website domain
github.com
Launched
June 19, 2026
Cohort
—
Upvotes
6
Upvotes percentile
0.31420765027322406
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hi everyone,I started working on nanoeuler after the ban of anthropic's fable because my ambition and dream is to work in the AI field in anthropic. The two interesting reasons that led me to create nanoeuler were the first, interfacing with llm does not mean understanding how they are composed and two, working on llm with a very low-level layer to understand the correlation between parameters and data and growth of the model and how the GPU works and how some layers can be optimized. So I started working on it with a research aspect by making nanoeuler grow more and more but doing one step after another starting from Shakespeare.txt and understanding what a text generation model understands at 23 million parameters. For example, nanoeuler at that number had understood that Name: started a line and wrote that line with sense. I wrote everything in CUDA because I wanted to not use any intermediary between the model in training and inference and what it had to do. Then the use of SFT and much more, even if in small ways, were really useful to understand the various step to make an llm like a chatbot.Any feedback, help, or suggestions are absolutely welcome!

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
gpt-2 scale model implementation in c and cuda
Manually corrected
False

Could you build this?

No Writing a GPT-2 architecture and training loop from scratch in pure C and custom CUDA kernels requires low-level GPU programming, memory management, and CUDA optimization expertise that AI assistants cannot reliably write or debug end-to-end.

What it would actually take: The stack involves pure C, CUDA C++, and direct interaction with NVIDIA GPU hardware (handling warp divergence, shared memory tiling, coalesced memory access, and custom GEMM/attention kernels). The hardest parts are implementing numerical stability, backpropagation gradient calculations across all layers, and profiling/optimizing kernel throughput with NVIDIA Nsight. This requires deep systems and high-performance computing (HPC) engineering experience.

Discussion

3 comments analyzed.

Competitors

Other products that read as similar to this one — 91 launches clear the similarity bar, closest 8 shown.

Attention rank: #64 of 92 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 231 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.