How I topped the HuggingFace open LLM leaderboard on two gaming GPUs
Details
- External ID
- 47322887
- Source
- HN
- Company
- —
- Product
- How I topped the HuggingFace open LLM leaderboard on two gaming GPUs
- Website domain
- github.io
- Launched
- March 10, 2026
- Cohort
- —
- Upvotes
- 495
- Upvotes percentile
- 0.995079950799508
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I found that duplicating a specific block of 7 middle layers in Qwen2-72B, without modifying any weights, improved performance across all Open LLM Leaderboard benchmarks and took #1. As of 2026, the top 4 models on that leaderboard are still descendants.The weird finding: single-layer duplication does nothing. Too few layers, nothing. Too many, it gets worse. Only circuit-sized blocks of ~7 layers work. This suggests pretraining carves out discrete functional circuits in the layer stack that only work when preserved whole.The whole thing was developed on 2x RTX 4090s in my basement. I'm now running current models (GLM-4.7, Qwen3.5, MiniMax M2.5) on a dual GH200 rig (see my other post). Code and new models coming soon.Happy to answer questions.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- method for training llms efficiently on consumer gpus
- Manually corrected
- False
Could you build this?
No This project is a scientific discovery in LLM internal interpretability and architecture manipulation, requiring novel empirical AI research rather than software development.
What it would actually take: To reproduce this, one needs PyTorch, Hugging Face Transformers, and tooling for hidden-state probe inspection (custom interpretability pipelines) running on multi-GPU hardware. The challenge lies in designing probing experiments to identify which functional layers correspond to reasoning versus representation, stitching tensor layers, and validating benchmarks like IFEval, MATH, and MMLU-Pro across 72B-parameter models.
Discussion
20 comments analyzed.
Competitors mentioned: DeepSeek V3.2 models, RAG and ICL approaches, MoE architectures
Concerns raised: Dimension alignment problem in model merging across different training runs, Performance ceiling limited by training data geometry and text bandwidth, Layer order and configuration sensitivity - interference when combining certain blocks, Computational cost increases with layer repetition
Feature requests: Combine layer duplication with automated architecture search (Karpathy's autoresearch), Test variable repetition patterns within identified blocks (e.g., 1,2,3,3,4,5,6,6,7), Analyze cross-correlation between layer outputs as alternative to benchmark sweep, Standardized layer architecture for dynamic downloading and swapping
Competitors
Other products that read as similar to this one — 45 launches clear the similarity bar, closest 8 shown.
Attention rank: #2 of 46 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 125 days after the earliest competitor.
- Duplicate 3 layers in a 24B LLM, logical deduction .22→.76. No training · hn · 2026-03-18 · 265 upvotes · similarity 0.57
- OS Megakernel that match M5 Max Tok/w at 2x the Throughput on RTX 3090 · hn · 2026-04-08 · 6 upvotes · similarity 0.42
- llm-compute-allocation-modeling · github · 2026-09-24 · 11 upvotes · similarity 0.40
- FuseCells · hn · 2025-12-31 · 6 upvotes · similarity 0.40
- Pencil Puzzle Bench · hn · 2026-03-03 · 5 upvotes · similarity 0.40
- LLM Onestop · hn · 2025-11-10 · 8 upvotes · similarity 0.38
- Reversing YouTube’s “Most Replayed” Graph · hn · 2026-01-16 · 87 upvotes · similarity 0.37
- bonsai2-small-gpu · github · 2026-09-19 · 54 upvotes · similarity 0.37
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.