Run TRELLIS.2 Image-to-3D generation natively on Apple Silicon
Details
- External ID
- 47828896
- Source
- HN
- Company
- —
- Product
- Run TRELLIS.2 Image-to-3D generation natively on Apple Silicon
- Website domain
- github.com
- Launched
- April 20, 2026
- Cohort
- —
- Upvotes
- 202
- Upvotes percentile
- 0.9627249357326478
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I ported Microsoft's TRELLIS.2 (4B parameter image-to-3D model) to run on Apple Silicon via PyTorch MPS. The original requires CUDA with flash_attn, nvdiffrast, and custom sparse convolution kernels: none of which work on Mac.I replaced the CUDA-specific ops with pure-PyTorch alternatives: a gather-scatter sparse 3D convolution, SDPA attention for sparse transformers, and a Python-based mesh extraction replacing CUDA hashmap operations. Total changes are a few hundred lines across 9 files.Generates ~400K vertex meshes from single photos in about 3.5 minutes on M4 Pro (24GB). Not as fast as H100 (where it takes seconds), but it works offline with no cloud dependency.https://github.com/shivampkumar/trellis-mac
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- image-to-3d generation on apple silicon
- Manually corrected
- False
Could you build this?
No Porting complex CUDA custom kernels, sparse convolutions, and differentiable rasterizers to pure PyTorch/MPS on Apple Silicon requires specialized low-level graphics and GPU kernel programming skills.
What it would actually take: The port requires reverse-engineering CUDA-specific implementations of nvdiffrast and flash_attn into Apple Metal Shading Language (MSL) or vectorized PyTorch tensor operations compatible with Metal Performance Shaders (MPS). The primary difficulty is handling sparse 3D convolutional operations and custom rasterizers without native CUDA acceleration while avoiding memory leaks on unified memory architecture. This requires advanced GPU systems programmers with expertise in both CUDA and Apple's Metal/MPS runtime.
Discussion
20 comments analyzed.
Competitors mentioned: ML Sharp (Apple's 3D Gaussian scene representation), Meshy (mesh generation tool), mtlgemm, mtldiffrast (alternative implementations)
Concerns raised: 16GB memory insufficient for fp32 models plus activations, TRELLIS.2 limited to single-image input only, Meshy has poor UI/UX (gamified, opaque, sleazy)
Feature requests: Multi-view input support, Sequential model loading instead of parallel for lower VRAM, Rewrite sparse convolutions in Metal shaders for faster performance, Standalone toolkit/writeup for sparse 3D convolution and SDPA attention swaps
Competitors
Other products that read as similar to this one — 160 launches clear the similarity bar, closest 8 shown.
Attention rank: #10 of 161 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 173 days after the earliest competitor.
- localmesh-engine · github · 2026-09-10 · 29 upvotes · similarity 0.47
- Apple's SHARP running in the browser via ONNX runtime web · hn · 2026-05-03 · 185 upvotes · similarity 0.46
- RunMat · hn · 2025-12-02 · 21 upvotes · similarity 0.45
- mvcc · github · 2026-09-23 · 51 upvotes · similarity 0.45
- macuda · github · 2026-09-16 · 30 upvotes · similarity 0.44
- code-map · github · 2026-09-17 · 48 upvotes · similarity 0.41
- Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac · hn · 2026-07-29 · 919 upvotes · similarity 0.40
- Magenta Real-Time Music Generation Locally on iPhone, Without the GPU · hn · 2026-06-10 · 9 upvotes · similarity 0.40
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.