Revibing nanochat's inference model in C++ with ggml
Details
- External ID
- 46566128
- Source
- HN
- Company
- —
- Product
- Revibing nanochat's inference model in C++ with ggml
- Website domain
- github.com
- Launched
- Jan. 10, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.09617918313570488
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Recently I wanted to see if I could vibe some serious C++ code.The result is a C++ re-implementation of Andrej Karpathy's nanochat's inferencing part (https://github.com/karpathy/nanochat), built on top of ggml. Unlike llama.cpp, this isn't a standalone binary; it is a C++ library & Python wrapper designed to swap out some core classes within the nanochat pipeline. For playability, I’ve tried to keep the dependencies to a minimum: just ggml, nanobind, and gtest for unit tests.Features and limitations:- A drop-in replacement of nanochat’s `GPT` and `KVCache` classes. So far I’ve only tested this with `chat_web.py`. You can see how it's integrated here: https://github.com/k-ye/nanochat/pull/1- Supports CPU and GPU (Metal yes, CUDA probably?).- Handles PyTorch-to-GGUF conversion automatically on demand.- Only float32 is currently supported.Benchmark:On an M3 Max (Metal), throughput is roughly 1/3 that of the original PyTorch implementation. I haven’t profiled the code yet, but I suspect the bottleneck is the lack of bf16 support.Motivation- Writing meaningful (& fun) C++ again: I used to spend a lot of my day-to-day time in C++ while working at various tech companies. These days, opportunities to use it for personal projects are rare, as it’s often hard to find a use case where C++'s advantages truly matter.- Testing "Vibe Coding" capabilities: Most of my current work is in UE5. Ironically, Blueprints—which were designed to help non-coders—have become a bottleneck in the LLM era... Admittedly, the AI agent has generated some FOMO in me, and I wanted to see if AI could handle a lower-level C++ implementation of a complex system from scratch.- Understanding the LLM internals.Why nanochat?It hits the "Goldilocks" zone: popular enough to be relevant, concise enough to be educational, and practical enough to deserve a serious C++ implementation.If you’re like me — an infra guy from the old days who feels a bit threatened by LLM and/or AI coding — I think nanochat is a great reference. Tinkering with it however you like is a nice way to demystify the tech. I relied heavily on Claude Code (CC) for the implementation. Overall, I am both impressed and genuinely pleased with the experience.Happy to answer questions, hear feedback or further discuss AI coding!
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- inference model optimization for developers
- Manually corrected
- False
Could you build this?
No Re-implementing a neural network inference engine and compute graph in C++ using low-level tensor libraries (ggml) requires deep systems programming and ML computational graph expertise.
What it would actually take: Requires deep familiarity with C/C++, ggml tensor operations, memory management, and transformer architectures (KV cache, RoPE, attention mechanisms). The developer must accurately map PyTorch/Python tensor ops into ggml C compute graphs without memory leaks or numerical discrepancies, then bind them to Python via pybind11/nanobind. AI assistants struggle significantly with low-level C memory models and debugging tensor shape/stride mismatches in custom C graph kernels.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 47 launches clear the similarity bar, closest 8 shown.
Attention rank: #47 of 48 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 61 days after the earliest competitor.
- NanoEuler · hn · 2026-06-28 · 55 upvotes · similarity 0.46
- NanoEuler · hn · 2026-06-19 · 6 upvotes · similarity 0.46
- GPT‑5.4 mini and nano · ph · 2026-03-18 · 267 upvotes · similarity 0.43
- Run 500B+ Parameter LLMs Locally on a Mac Mini · hn · 2026-03-09 · 17 upvotes · similarity 0.40
- Llmtop · hn · 2026-03-18 · 5 upvotes · similarity 0.40
- Tiny-vLLM · hn · 2026-05-29 · 205 upvotes · similarity 0.39
- OpenGraviton · hn · 2026-03-07 · 13 upvotes · similarity 0.39
- Sipp · hn · 2026-06-24 · 5 upvotes · similarity 0.38
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.