Gemma 4 Multimodal Fine-Tuner for Apple Silicon
Details
- External ID
- 47680309
- Source
- HN
- Company
- —
- Product
- Gemma 4 Multimodal Fine-Tuner for Apple Silicon
- Website domain
- github.com
- Launched
- April 7, 2026
- Cohort
- —
- Upvotes
- 235
- Upvotes percentile
- 0.9665809768637532
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
About six months ago, I started working on a project to fine-tune Whisper locally on my M2 Ultra Mac Studio with a limited compute budget. I got into it. The problem I had at the time was I had 15,000 hours of audio data in Google Cloud Storage, and there was no way I could fit all the audio onto my local machine, so I built a system to stream data from my GCS to my machine during training.Gemma 3n came out, so I added that. Kinda went nuts, tbh.Then I put it on the shelf.When Gemma 4 came out a few days ago, I dusted it off, cleaned it up, broke out the Gemma part from the Whisper fine-tuning and added support for Gemma 4.I'm presenting it for you here today to play with, fork and improve upon.One thing I have learned so far: It's very easy to OOM when you fine-tune on longer sequences! My local Mac Studio has 64GB RAM, so I run out of memory constantly.Anywho, given how much interest there is in Gemma 4, and frankly, the fact that you can't really do audio fine-tuning with MLX, that's really the reason this exists (in addition to my personal interest). I would have preferred to use MLX and not have had to make this, but here we are. Welcome to my little side quest.And so I made this. I hope you have as much fun using it as I had fun making it.-Matt
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- gemma 4 fine-tuner for apple silicon
- Manually corrected
- False
Could you build this?
Partial While Python fine-tuning harnesses exist, building an Apple Silicon (MLX/Metal) pipeline that streams 15,000 hours of multimodal training data without filling local storage requires complex memory and I/O engineering.
What it would actually take: A functioning version uses Apple's MLX framework or PyTorch MPS with custom LoRA/QLoRA kernels for Gemma multimodal architectures. The challenging piece is building a high-throughput streaming dataloader that dynamically chunks, decodes, and caches massive audio streams directly from cloud storage into Apple unified memory without hitting disk exhaustion or crashing the unified memory pool. This requires deep familiarity with low-level audio DSP, Apple Silicon memory bandwidth management, and distributed streaming dataloaders.
Discussion
20 comments analyzed.
Competitors mentioned: Whisper v3 large, Parakeet / parakeet-mlx, MacWhisper, Gemini / Gemini Pro, ChatMCP
Concerns raised: Security vulnerabilities in GGUF and pickle formats (CVE-2024-34359 RCE), Memory usage explodes when processing audio beyond 30 seconds, Smaller models have poor quality despite faster inference, Cannot fine-tune on Apple Silicon, Whisper has 30-second context window limitation
Feature requests: Fine-tune models for specific accents locally, Better batch inference instead of linear processing, Distill large models (Gemini) into tiny on-device models, VAD preprocessing to filter junk audio data, Support for longer audio sequences without memory explosion
Competitors
Other products that read as similar to this one — 158 launches clear the similarity bar, closest 8 shown.
Attention rank: #12 of 159 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 151 days after the earliest competitor.
- TongueType · hn · 2026-05-15 · 5 upvotes · similarity 0.50
- Ableton Live MCP · hn · 2026-05-03 · 123 upvotes · similarity 0.48
- Google Gemma 4 12B · ph · 2026-06-04 · 309 upvotes · similarity 0.48
- Cactus Hybrid: We taught Gemma 4 to know when it's wrong · hn · 2026-07-22 · 191 upvotes · similarity 0.48
- Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac · hn · 2026-07-29 · 919 upvotes · similarity 0.46
- SoundFerry · ph · 2026-09-26 · 1 upvotes · similarity 0.44
- Writekin · hn · 2026-07-28 · 6 upvotes · similarity 0.44
- Audio Player with "Binaural Beats" tuned to the same key as your music · hn · 2026-07-24 · 21 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.