Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Text-to-video model from scratch (2 brothers, 2 years, 2B params)

Details

External ID
46721488
Source
HN
Company
—
Product
Text-to-video model from scratch (2 brothers, 2 years, 2B params)
Website domain
huggingface.co
Launched
Jan. 22, 2026
Cohort
—
Upvotes
158
Upvotes percentile
0.9446640316205533
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Writeup (includes good/bad sample generations): https://www.linum.ai/field-notes/launch-linum-v2We're Sahil and Manu, two brothers who spent the last 2 years training text-to-video models from scratch. Today we're releasing them under Apache 2.0.These are 2B param models capable of generating 2-5 seconds of footage at either 360p or 720p. In terms of model size, the closest comparison is Alibaba's Wan 2.1 1.3B. From our testing, we get significantly better motion capture and aesthetics.We're not claiming to have reached the frontier. For us, this is a stepping stone towards SOTA - proof we can train these models end-to-end ourselves.Why train a model from scratch?We shipped our first model in January 2024 (pre-Sora) as a 180p, 1-second GIF bot, bootstrapped off Stable Diffusion XL. Image VAEs don't understand temporal coherence, and without the original training data, you can't smoothly transition between image and video distributions. At some point you're better off starting over.For v2, we use T5 for text encoding, Wan 2.1 VAE for compression, and a DiT-variant backbone trained with flow matching. We built our own temporal VAE but Wan's was smaller with equivalent performance, so we used it to save on embedding costs. (We'll open-source our VAE shortly.)The bulk of development time went into building curation pipelines that actually work (e.g., hand-labeling aesthetic properties and fine-tuning VLMs to filter at scale).What works: Cartoon/animated styles, food and nature scenes, simple character motion. What doesn't: Complex physics, fast motion (e.g., gymnastics, dancing), consistent text.Why build this when Veo/Sora exist? Products are extensions of the underlying model's capabilities. If users want a feature the model doesn't support (character consistency, camera controls, editing, style mapping, etc.), you're stuck. To build the product we want, we need to update the model itself. That means owning the development process. It's a bet that will take time (and a lot of GPU compute) to pay off, but we think it's the right one.What’s next? - Post-training for physics/deformations - Distillation for speed - Audio capabilities - Model scalingWe kept a “lab notebook” of all our experiments in Notion. Happy to answer questions about building a model from 0 → 1. Comments and feedback welcome!

Enrichment

Theme
ai video creation and repurposing
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
text-to-video model
Manually corrected
False

Could you build this?

No Training a 2B-parameter text-to-video foundation model from scratch requires advanced ML research expertise, massive GPU compute clusters, and custom distributed training pipelines.

What it would actually take: Building Linum v2 required curating and filtering hundreds of millions of video-text pairs, building a 3D spatio-temporal diffusion or DiT architecture in PyTorch, and managing distributed multi-node training clusters with FlashAttention, DeepSpeed, or Megatron-LM. The primary barriers are deep generative video modeling research experience and tens or hundreds of thousands of dollars in high-end compute (e.g., H100s).

Discussion

20 comments analyzed.

Competitors mentioned: Project #5 at bytebyteai.com, Karpathy's courses (future)

Concerns raised: High VRAM requirements (20GB+) limits consumer hardware accessibility, T5 text encoder (5B parameters) disproportionately large for 2B video model, Memory footprint makes inference economics unviable for consumer hardware, Substantial compute cost to train

Feature requests: Optimize/quantize T5 encoder to 8-bit or 4-bit, Add option to delete T5 from memory after text encoding, Manual layer offloading to reduce VRAM usage, End-to-end guide/course for building video models, Public lab notebook with experiments

Competitors

Other products that read as similar to this one — 136 launches clear the similarity bar, closest 8 shown.

Attention rank: #10 of 137 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 85 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.