I embedded 685M public texts in 32 minutes (on 8x A100, Rust, TensorRT)
Details
- External ID
- 48400053
- Source
- HN
- Company
- —
- Product
- I embedded 685M public texts in 32 minutes (on 8x A100, Rust, TensorRT)
- Website domain
- github.com
- Launched
- June 4, 2026
- Cohort
- —
- Upvotes
- 7
- Upvotes percentile
- 0.42008196721311475
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Quick note on how it works and how I've done my batch embedding engine IgniteMS.The whole thing runs as one process using Rust, reading input, tokenizing, packing batches, keeping the queue full. TensorRT handles inference. Python is only as a wrapper.I built it this way because when you use more than couple of GPUs, the GPUs stop being the problem. CPU cannot feed them fast enough. One A100 can go through batches faster than Python can tokenize and feed, so the GPU just sits there idle waiting for work. Most of my time went into optimizing this. At 8 GPUs that was basically the entire challenge.On cost. I ran the big 2B messages job on a spot p4d instance (8x A100 40GB). After filtering and dedupping I got 685M raw texts. With my new engine the whole production run finishes in about half an hour. Previously I used on-demand for these jobs, now switched to spots. If AWS reclaims the box, I just rerun it. It's roughly $7 for half-an-hour run. And at least right now spots are easier to get than on-demand.Open warning: it's batch only and NVIDIA only. You can use it both as a docker image and native. I used some optimizations for my production run. With default settings you can expect to see ~250K msg/sec if you run the benchmark script on your p4d box. https://github.com/Artain-AI/ignite-ms/blob/main/BENCHMARKIN...v1.1.0 added TensorRT 11 and 60 models, 23 tested on 1x and 4x A100.Happy to share details.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Data infrastructure
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- text embedding infrastructure
- Manually corrected
- False
Could you build this?
No Achieving throughput of 685M texts in 32 minutes across multi-GPU setups demands expert-level systems programming in Rust and deep familiarity with TensorRT and CUDA memory pipelines.
What it would actually take: The architecture requires high-concurrency Rust with zero-copy ring buffers, custom HuggingFace tokenization pipelines saturating multiple CPU cores, and pinned host memory transfers. Inference is backed by TensorRT engines running across an 8x A100 NVLink cluster with dynamic batch packing and asynchronous CUDA streams to prevent pipeline stalls.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 138 launches clear the similarity bar, closest 8 shown.
Attention rank: #86 of 139 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 218 days after the earliest competitor.
- Optimizing LiteLLM with Rust · hn · 2025-11-18 · 27 upvotes · similarity 0.48
- RunMat · hn · 2025-12-02 · 21 upvotes · similarity 0.47
- I built Wool, a lightweight distributed Python runtime · hn · 2026-03-14 · 15 upvotes · similarity 0.46
- Sub-millisecond VM sandboxes using CoW memory forking · hn · 2026-03-17 · 311 upvotes · similarity 0.44
- Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU · hn · 2026-06-30 · 10 upvotes · similarity 0.43
- Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload · hn · 2026-07-26 · 22 upvotes · similarity 0.43
- I made an open-source Rust program for memory-efficient genomics · hn · 2025-11-13 · 17 upvotes · similarity 0.42
- seedance2.5-seedance-2.5-prompts · github · 2026-09-24 · 53 upvotes · similarity 0.40
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a data infrastructure tool for Media & entertainment yet.