Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Gemma 4 Multimodal Fine-Tuner for Apple Silicon

Details

External ID
47680309
Source
HN
Company
—
Product
Gemma 4 Multimodal Fine-Tuner for Apple Silicon
Website domain
github.com
Launched
April 7, 2026
Cohort
—
Upvotes
235
Upvotes percentile
0.9665809768637532
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

About six months ago, I started working on a project to fine-tune Whisper locally on my M2 Ultra Mac Studio with a limited compute budget. I got into it. The problem I had at the time was I had 15,000 hours of audio data in Google Cloud Storage, and there was no way I could fit all the audio onto my local machine, so I built a system to stream data from my GCS to my machine during training.Gemma 3n came out, so I added that. Kinda went nuts, tbh.Then I put it on the shelf.When Gemma 4 came out a few days ago, I dusted it off, cleaned it up, broke out the Gemma part from the Whisper fine-tuning and added support for Gemma 4.I'm presenting it for you here today to play with, fork and improve upon.One thing I have learned so far: It's very easy to OOM when you fine-tune on longer sequences! My local Mac Studio has 64GB RAM, so I run out of memory constantly.Anywho, given how much interest there is in Gemma 4, and frankly, the fact that you can't really do audio fine-tuning with MLX, that's really the reason this exists (in addition to my personal interest). I would have preferred to use MLX and not have had to make this, but here we are. Welcome to my little side quest.And so I made this. I hope you have as much fun using it as I had fun making it.-Matt

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
gemma 4 fine-tuner for apple silicon
Manually corrected
False

Could you build this?

Partial While Python fine-tuning harnesses exist, building an Apple Silicon (MLX/Metal) pipeline that streams 15,000 hours of multimodal training data without filling local storage requires complex memory and I/O engineering.

What it would actually take: A functioning version uses Apple's MLX framework or PyTorch MPS with custom LoRA/QLoRA kernels for Gemma multimodal architectures. The challenging piece is building a high-throughput streaming dataloader that dynamically chunks, decodes, and caches massive audio streams directly from cloud storage into Apple unified memory without hitting disk exhaustion or crashing the unified memory pool. This requires deep familiarity with low-level audio DSP, Apple Silicon memory bandwidth management, and distributed streaming dataloaders.

Discussion

20 comments analyzed.

Competitors mentioned: Whisper v3 large, Parakeet / parakeet-mlx, MacWhisper, Gemini / Gemini Pro, ChatMCP

Concerns raised: Security vulnerabilities in GGUF and pickle formats (CVE-2024-34359 RCE), Memory usage explodes when processing audio beyond 30 seconds, Smaller models have poor quality despite faster inference, Cannot fine-tune on Apple Silicon, Whisper has 30-second context window limitation

Feature requests: Fine-tune models for specific accents locally, Better batch inference instead of linear processing, Distill large models (Gemini) into tiny on-device models, VAD preprocessing to filter junk audio data, Support for longer audio sequences without memory explosion

Competitors

Other products that read as similar to this one — 158 launches clear the similarity bar, closest 8 shown.

Attention rank: #12 of 159 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 151 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.