Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Run TRELLIS.2 Image-to-3D generation natively on Apple Silicon

Details

External ID
47828896
Source
HN
Company
—
Product
Run TRELLIS.2 Image-to-3D generation natively on Apple Silicon
Website domain
github.com
Launched
April 20, 2026
Cohort
—
Upvotes
202
Upvotes percentile
0.9627249357326478
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

I ported Microsoft's TRELLIS.2 (4B parameter image-to-3D model) to run on Apple Silicon via PyTorch MPS. The original requires CUDA with flash_attn, nvdiffrast, and custom sparse convolution kernels: none of which work on Mac.I replaced the CUDA-specific ops with pure-PyTorch alternatives: a gather-scatter sparse 3D convolution, SDPA attention for sparse transformers, and a Python-based mesh extraction replacing CUDA hashmap operations. Total changes are a few hundred lines across 9 files.Generates ~400K vertex meshes from single photos in about 3.5 minutes on M4 Pro (24GB). Not as fast as H100 (where it takes seconds), but it works offline with no cloud dependency.https://github.com/shivampkumar/trellis-mac

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
image-to-3d generation on apple silicon
Manually corrected
False

Could you build this?

No Porting complex CUDA custom kernels, sparse convolutions, and differentiable rasterizers to pure PyTorch/MPS on Apple Silicon requires specialized low-level graphics and GPU kernel programming skills.

What it would actually take: The port requires reverse-engineering CUDA-specific implementations of nvdiffrast and flash_attn into Apple Metal Shading Language (MSL) or vectorized PyTorch tensor operations compatible with Metal Performance Shaders (MPS). The primary difficulty is handling sparse 3D convolutional operations and custom rasterizers without native CUDA acceleration while avoiding memory leaks on unified memory architecture. This requires advanced GPU systems programmers with expertise in both CUDA and Apple's Metal/MPS runtime.

Discussion

20 comments analyzed.

Competitors mentioned: ML Sharp (Apple's 3D Gaussian scene representation), Meshy (mesh generation tool), mtlgemm, mtldiffrast (alternative implementations)

Concerns raised: 16GB memory insufficient for fp32 models plus activations, TRELLIS.2 limited to single-image input only, Meshy has poor UI/UX (gamified, opaque, sleazy)

Feature requests: Multi-view input support, Sequential model loading instead of parallel for lower VRAM, Rewrite sparse convolutions in Metal shaders for faster performance, Standalone toolkit/writeup for sparse 3D convolution and SDPA attention swaps

Competitors

Other products that read as similar to this one — 160 launches clear the similarity bar, closest 8 shown.

Attention rank: #10 of 161 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 173 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.