Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Litelink

local-first, embedded stream capture into Iceberg tables

Details

External ID
49549760
Source
HN
Company
—
Product
Litelink
Website domain
github.com
Launched
Sept. 3, 2026
Cohort
—
Upvotes
9
Upvotes percentile
0.5422647527910686
Tags
—
Fetched at
Sept. 10, 2026, 5:31 a.m.
Updated at
Sept. 10, 2026, 5:31 a.m.

Description

Hi HN! I just wanted to share litelink a local-first, embedded capture library I built in python (code is heavily AI generated but designed and reviewed by yours truly). I've been using this for point-and-shoot WebSocket capture but I imagine it could also be useful for observability/metrics ingestion as well. Litelink supports a single writer per stream.I've been doing a lot of development and deployments on tiny VMs (2 vCPU, 8GB, 50-100GB disk) and didn't want the complexity or cost of managing central brokers (Kafka), databases (Postgres), and CDC/connectors just to get queryable WebSocket stream capture running.With litelink, you configure a log in code, and end-to-end setup takes <5 minutes (see the example scripts in the repo). The log is itself an Iceberg table (actually two: a local and archive table), so there's no second copy of your data to keep in sync or connector to manage.I'm sure there are still bugs, but I recently migrated all the capture feeds for a personal research project to litelink, and the experience has been night and day. Before that, I'd hand-rolled a capture system and was dealing with all the issues you'd expect (e.g. small file problem). I'll post some before/after stats in a comment below.I tried to channel the same ethos as LanceDB/Iceberg/SQLite. Everything runs local first without a network connection required. I've tried to abstract the complexity of stream/data lifecycle maintenance away behind a few public library methods. Hopefully someone else finds this useful! Let me know what you think.repo: https://github.com/nhobin219/litelinkspec: https://github.com/nhobin219/litelink/blob/main/docs/SPEC.mdpypi: `pip install litelink`

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Data infrastructure
Audience
Developer
AI stance
Not AI
Project type
Commercial product
Normalized one-liner
local-first stream capture to iceberg tables
Manually corrected
False

Could you build this?

Partial A Python capture script is simple, but reliably streaming high-throughput events into Apache Iceberg tables locally without heavy JVM overhead requires low-level Parquet writing and metadata catalog management.

What it would actually take: Uses PyIceberg or PyArrow with DuckDB to serialize streaming WebSocket JSON frames directly into Apache Parquet format while handling Iceberg schema evolution, snapshot generation, and manifest list commits. The hard part is managing micro-batching and transactional Iceberg metadata ACID guarantees without running external heavy distributed services or memory leaks under continuous ingestion. Requires data engineering expertise in storage formats and tabular streaming specifications.

Discussion

2 comments analyzed.

Concerns raised: Litelink overhead dominates in small datasets due to auto-generated offset column, Storage efficiency worse than legacy system for certain data types, File count explosion compared to legacy approach

Feature requests: Optimize or eliminate litelink_offset column overhead for small rows, Compression improvements for low-compression datasets

Competitors

Other products that read as similar to this one — 54 launches clear the similarity bar, closest 8 shown.

Attention rank: #34 of 55 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 306 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a data infrastructure tool for Media & entertainment yet.