Lance
image/video generation and understanding in one model
Details
- External ID
- 48209668
- Source
- HN
- Company
- —
- Product
- Lance
- Website domain
- github.com
- Launched
- May 20, 2026
- Cohort
- —
- Upvotes
- 64
- Upvotes percentile
- 0.8618739903069467
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
The model has 3B active parameters. We put the code, homepage, paper and model links here:- Code: https://github.com/bytedance/Lance- Homepage: https://lance-project.github.io/- Paper: https://arxiv.org/abs/2605.18678- Model: https://huggingface.co/bytedance-research/Lancep.s. Lance is a research project, not a polished product. The model was trained using fewer than 128 GPUs.
Enrichment
- Theme
- code-driven AI video and animation
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- image and video generation and understanding model
- Manually corrected
- False
Could you build this?
No Lance is a multimodal 3B foundation model for combined image/video understanding and generation, requiring massive compute clusters, custom training pipelines, and deep machine learning research expertise.
What it would actually take: Developing Lance requires a unified autoregressive or diffusion transformer architecture capable of processing interleaved video, image, and text tokens. The engineering pipeline requires distributed training across large GPU clusters using Megatron/DeepSpeed, petabyte-scale curated video-image-text datasets, and novel loss formulations for joint generative-discriminative multimodal representations.
Discussion
15 comments analyzed.
Competitors mentioned: vLLM, sglang, lancedb
Concerns raised: Video output resolution and frame rate are low (720p), Samples upscaled and frame-interpolated, masking actual capabilities, Current agents struggle with unconventional UI screenshots, Requires 40GB VRAM despite 3B active parameters
Feature requests: Port to sglang or vLLM, Better performance on navigating and recording actual applications/UX
Competitors
Other products that read as similar to this one — 20 launches clear the similarity bar, closest 8 shown.
Attention rank: #5 of 21 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 172 days after the earliest competitor.
- Text-to-video model from scratch (2 brothers, 2 years, 2B params) · hn · 2026-01-22 · 158 upvotes · similarity 0.42
- GPTProto · ph · 2026-09-26 · 1 upvotes · similarity 0.34
- motion-graphics-skills · github · 2026-09-26 · 30 upvotes · similarity 0.33
- High speed graphics rendering research with tinygrad/tinyJIT · hn · 2026-01-22 · 31 upvotes · similarity 0.33
- flux-3-dev.github.io · github · 2026-09-21 · 11 upvotes · similarity 0.33
- wan-3.run · ph · 2026-09-14 · 1 upvotes · similarity 0.33
- Step 3.7 Flash · ph · 2026-05-30 · 183 upvotes · similarity 0.32
- Lorem.video · hn · 2026-02-11 · 6 upvotes · similarity 0.32
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.