Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Duplicate 3 layers in a 24B LLM, logical deduction .22→.76. No training

Details

External ID
47431671
Source
HN
Company
—
Product
Duplicate 3 layers in a 24B LLM, logical deduction .22→.76. No training
Website domain
github.com
Launched
March 18, 2026
Cohort
—
Upvotes
265
Upvotes percentile
0.97539975399754
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

I replicated David Ng's RYS method (https://dnhkng.github.io/posts/rys/) on consumer AMD GPUs (RX 7900 XT + RX 6950 XT) and found something I didn't expect.Transformers appear to have discrete "reasoning circuits" — contiguous blocks of 3-4 layers that act as indivisible cognitive units. Duplicate the right block and the model runs its reasoning pipeline twice. No weights change. No training. The model just thinks longer.The results on standard benchmarks (lm-evaluation-harness, n=50):Devstral-24B, layers 12-14 duplicated once: - BBH Logical Deduction: 0.22 → 0.76 - GSM8K (strict): 0.48 → 0.64 - MBPP (code gen): 0.72 → 0.78 - Nothing degradedQwen2.5-Coder-32B, layers 7-9 duplicated once: - Reasoning probe: 76% → 94%The weird part: different duplication patterns create different cognitive "modes" from the same weights. Double-pass boosts math. Triple-pass boosts emotional reasoning. Interleaved doubling (13,13,14,14,15,15,16) creates a pure math specialist. Same model, same VRAM, different routing.The circuit boundaries are sharp — shift by one layer and the effect disappears or inverts. Smaller models (24B) have tighter circuits (3 layers) than larger ones (Ng found 7 layers in 72B).Tools to find circuits in any GGUF model and apply arbitrary layer routing are in the repo. The whole thing — sweep, discovery, validation — took one evening.Happy to answer questions.

Enrichment

Theme
ML inference and model optimization
Vertical
—
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
llm layer optimization research
Manually corrected
False

Could you build this?

No This is exploratory deep learning / mechanistic interpretability research involving model architecture manipulation and AMD ROCm GPU compute experiments rather than app development.

What it would actually take: Replicating this requires PyTorch/ROCm deep learning environments, custom tensor manipulation scripts to split and splice HuggingFace transformer model weights/layers, and evaluation harnesses (such as lm-evaluation-harness for logic benchmarks like ARC or GSM8K). It requires specialized knowledge of LLM internals, layer architectures, KV-cache behavior, and deep learning runtime mechanics.

Discussion

20 comments analyzed.

Concerns raised: No statistically significant gains shown in results, Benefits collapse when including layers beyond optimal range, Circuit boundaries don't translate across different tasks, RLHF may degrade reasoning performance, Results show model lost performance on some categories

Feature requests: Train reasoning block separately on smaller corpus with fixed input/output translation blocks, Experiment with making reasoning block wider or narrower, Implement dynamic routing model instead of fixed layer selection, Add ability to replay/reconstruct intermediate layer outputs

Competitors

Other products that read as similar to this one — 70 launches clear the similarity bar, closest 8 shown.

Attention rank: #6 of 71 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 130 days after the earliest competitor.

Other launches for this product