Mini-AGI
Dynamic continual learning model trained on 8GB VRAM
Details
- External ID
- 49783133
- Source
- HN
- Company
- —
- Product
- mini-AGI
- Website domain
- github.com
- Launched
- Sept. 21, 2026
- Cohort
- —
- Upvotes
- 277
- Upvotes percentile
- 0.9808612440191388
- Tags
- —
- Fetched at
- Sept. 25, 2026, 5:02 p.m.
- Updated at
- Sept. 25, 2026, 5:02 p.m.
Description
Sorry for the pretentious name, I know, I know.. It just contains all the pieces I would like to see a AGI model to have, and I can't stand the temptation. Before throwing rocks at me, please take a glance at the Readme, and I hope it will cover your mood a little bit.So, first of all it does work and you can see the sample from the whole training run here: https://raw.githubusercontent.com/volotat/mini-AGI/refs/head...Here is the scaling law graph I have so far, and it looks very promising: https://github.com/volotat/mini-AGI/blob/main/assets/scaling...The model was built under my deep dissatisfaction so we cannot really train even moderately big models (1B+ scale) on the consumer's hardware. We can inference and fine-tune them for sure, but I would like to have full control over what the model sees over the training run, so it is fully aligned with my interests, not some corporations.I was thinking about for some time and come up with two interesting ideas I thought worth pursuing: MoE with a lot of experts that gets added and pruned from the model while it trains, where only a small subset of of experts are actually in use at any particular moment + batch 1 training on the single continuous stream of data.First allows us to be bounded only by the disk space in terms of number of parameters and load and unload experts only when they are needed. The second (if figured out and it turns out to be doable) allows us to get aways with small VRAM capacity because we do not need to store big randomized batches and their respective gradients.I started brainstorming with Claude and after some time we found an approach that seems to be promising, and low and behold, a few weeks pass and you can see the results yourself.Obviously, I did use AI in the process of making this project and I am pretty sure it would be completely impossible for me to do something like this without it, so I hope it is more than justified.The model is still running over the first of 7.8B characters corpus I selected for training, so the weights are not out yet, and it's about a couple weeks of waiting until they are cooked at the current reading speed. And yeah, the model just read continuous interleaved passages from the dataset, each by 32K characters long each as a single stream. Just as you or I would do.The set up seems to be really simple so you can git clone the project, run it and observe everything for yourself.Thanks for your attention.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- continual learning model runnable on consumer gpus
- Manually corrected
- False
Could you build this?
No Designing and training a dynamic continual learning model from scratch on tight memory limits (8GB VRAM) requires novel machine learning research and low-level PyTorch/CUDA optimization to prevent catastrophic forgetting.
What it would actually take: This requires implementing novel continually learning neural network architectures (e.g., progressive neural networks, sparse weight allocation, or replay buffers optimized for streaming batch-1 data) with custom PyTorch or Triton kernels to run efficiently on 8GB consumer VRAM. The developer needs deep ML research expertise in catastrophic forgetting, online gradient updates, and memory-efficient training mechanics. Standard vibe-coding prompts fail here because the foundational model architecture and update rules cannot be imported from standard off-the-shelf APIs.
Discussion
20 comments analyzed.
Competitors mentioned: RWKV, OpenClaw, Hermes
Concerns raised: calling it AGI is inaccurate and misleading, generates nonsensical outputs, still susceptible to catastrophic forgetting, graphs lack axis labels and explanations
Feature requests: axis labels and explanations for graphs
Competitors
Other products that read as similar to this one — 147 launches clear the similarity bar, closest 8 shown.
Attention rank: #9 of 148 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 321 days after the earliest competitor.
- Model-agnostic cognitive architecture for LLMs · hn · 2025-11-18 · 6 upvotes · similarity 0.51
- RunNburn · hn · 2026-07-30 · 11 upvotes · similarity 0.47
- Serve 100 Large AI models on a single GPU with low impact to TTFT · hn · 2025-11-08 · 7 upvotes · similarity 0.46
- The Analog I · hn · 2026-01-16 · 29 upvotes · similarity 0.45
- OpenGraviton · hn · 2026-03-07 · 13 upvotes · similarity 0.45
- An unmetered LLM API–$6/month, no token tracking, no limits · hn · 2026-07-06 · 12 upvotes · similarity 0.44
- A new engine to run Kimi K3 on a laptop · hn · 2026-07-29 · 7 upvotes · similarity 0.44
- Run 500B+ Parameter LLMs Locally on a Mac Mini · hn · 2026-03-09 · 17 upvotes · similarity 0.43
Other launches for this product
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.