NanoRL
RL training for LLMs in ~1,800 lines
Details
- External ID
- 49286216
- Source
- HN
- Company
- —
- Product
- NanoRL
- Website domain
- github.com
- Launched
- Aug. 13, 2026
- Cohort
- —
- Upvotes
- 11
- Upvotes percentile
- 0.6088709677419355
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:32 a.m.
- Updated at
- Sept. 10, 2026, 5:32 a.m.
Description
The smallest async RL trainer I could write: one loop that runs REINFORCE on CartPole on a laptop and async GRPO on a cluster (e.g. 8xH100 trainer, 8 vLLM workers, ran as a [SkyPilot job group](https://docs.skypilot.ai/en/latest/examples/job-groups.html) on k8s ).All without Ray or TRL or DeepSpeed etc., workers talk to the trainer over stdlib HTTP.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- reinforcement learning training framework for language models
- Manually corrected
- False
Could you build this?
No Implementing a distributed asynchronous RL trainer (GRPO/REINFORCE) that orchestrates GPU clusters and inference engines without existing frameworks requires deep expertise in distributed ML systems and RL mathematics.
What it would actually take: The stack involves PyTorch, custom CUDA/C++ extensions or optimized Triton kernels, raw socket/NCCL communication or custom async queues, and direct integration with vLLM worker processes. The hard part is managing distributed model weights synchronization, asynchronous off-policy gradient corrections, memory management during rollout generation, and low-latency cluster scheduling across H100 nodes without high-level abstractions like Ray or DeepSpeed.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 17 launches clear the similarity bar, closest 8 shown.
Attention rank: #11 of 18 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 285 days after the earliest competitor.
- I built a tiny LLM to demystify how language models work · hn · 2026-04-06 · 915 upvotes · similarity 0.38
- Mini-vLLM in ~500 lines of Python · hn · 2025-12-28 · 5 upvotes · similarity 0.37
- Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO) · hn · 2026-08-01 · 21 upvotes · similarity 0.35
- Run 500B+ Parameter LLMs Locally on a Mac Mini · hn · 2026-03-09 · 17 upvotes · similarity 0.35
- I RL-trained an agent that trains models with RL (for ~$1.3k) · hn · 2026-07-14 · 107 upvotes · similarity 0.34
- GLM Mini · ph · 2026-09-14 · 1 upvotes · similarity 0.33
- A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo) · hn · 2026-08-10 · 79 upvotes · similarity 0.33
- Micro-RLE ≤264-byte compression for UART/MCU logs, zero RAM growth · hn · 2025-11-01 · 7 upvotes · similarity 0.33
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.