Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Overfitted a 900KB Transformer to Compress a 100MB CSV into 7MB

Details

External ID
48644463
Source
HN
Company
—
Product
—
Website domain
—
Launched
June 23, 2026
Cohort
—
Upvotes
112
Upvotes percentile
0.9241803278688525
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

I built an experiment that uses an overfitted transformer and arithmetic coding to compress individual files.Instead of training the model to generalize, I train a 900KB transformer to memorize a single file and predict the next byte. Those predictions are fed into an arithmetic coder to produce the compressed output.On a 100MB NYC taxi CSV, it compresses to about 7MB (~0.5 bits/byte). On a 100MB slice of enwik9, it compresses to about 21MB (~1.68 bits/byte).It's pretty slow right now (roughly 20–30 minutes of training and 45 minutes each for compression and decompression on my AMD 7800XT).Checkout the repo - https://github.com/samyak112/pym-particles

Enrichment

Theme
file transfer and sharing tools
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
transformer-based csv compression
Manually corrected
False

Could you build this?

No This is a specialized machine learning and information theory research experiment involving custom miniature transformer architectures, next-byte entropy modeling, and integration with an arithmetic coding pipeline.

What it would actually take: The architecture involves PyTorch/JAX training an extremely compact auto-regressive transformer model to minimize cross-entropy loss on a single dataset, paired with an arithmetic coding implementation (often in C++ or Rust for performance). The hard parts are tuning model capacity vs. parameter footprint to avoid under/over-saturating the arithmetic coder, and handling finite-precision arithmetic coder stability with neural network logits. It requires deep research expertise in data compression, entropy coding, and neural network optimization.

Discussion

20 comments analyzed.

Competitors mentioned: ZPAQ compression algorithm, TabPFN v2, Conventional compression algorithms (Matt Mahoney's benchmarks), TensorDyne GPU chip for neural operations

Concerns raised: Model only works for specific file; retraining required for each new file, How to securely share the model beforehand for decryption without RSA, Very slow decompression (45 minutes for 100MB), Accuracy concerns with overfitted transformer adding random bytes, Heavy computational load running transformer in user's browser

Feature requests: Support for multiple files without retraining, Faster decompression speed, Better documentation of prior work and related approaches

Competitors

Other products that read as similar to this one — 139 launches clear the similarity bar, closest 8 shown.

Attention rank: #6 of 140 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 234 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.