Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

500k+ events/sec transformations for ClickHouse ingestion

Details

External ID
47693407
Source
HN
Company
—
Product
500k+ events/sec transformations for ClickHouse ingestion
Website domain
github.com
Launched
April 8, 2026
Cohort
—
Upvotes
13
Upvotes percentile
0.6658097686375322
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hi HN! We are Ashish and Armend, founders of GlassFlow.Over the last year, we worked with teams running high-throughput pipelines into self-hosted ClickHouse. Mostly for observability and real-time analytics.A question that came repeatedly was: What happens when throughput grows?Usually, things work fine at 10k events/sec, but we started seeing backpressure and errors at >100k.When the throughput per pipeline stops scaling, then adding more CPU/memory doesn’t help because often parts of the pipeline are not parallelized or are bottlenecked by state handling.At this point, engineers usually scale by adding more pipeline instances.That works but comes with some trade-offs: - You have to split the workload (e.g., multiple pipelines reading from the same source) - Transformation logic gets duplicated across pipelines - Stateful logic becomes harder to manage and keep consistent - Debugging and changes get more difficult because the data flow is fragmentedAnother challenge arises when working with high-cardinality keys like user IDs, session IDs, or request IDs, and when you need to handle longer time windows (24h or more). The state grows quickly and many systems rely on in-memory state, which makes it expensive and harder to recover from failures.We wanted to solve this problem and rebuild our approach at GlassFlow.Instead of scaling by adding more pipelines, we scale within a single pipeline by using replicas. Each replica consumes, processes, and writes independently, and the workload is distributed across them.In the benchmarks we’re sharing, this scales to 500k+ events/sec while still running stateful transformations and writing into ClickHouse.A few things we think are interesting: - Scaling is close to linear as you add replicas - Works with stateful transformations (not just stateless ingestion) - State is backed by a file-based KV store instead of relying purely on memory - The ClickHouse sink is optimized for batching to avoid small inserts - The product is built with GoFull write-up + benchmarks: https://www.glassflow.dev/blog/glassflow-now-scales-to-500k-...Repo: https://github.com/glassflow/clickhouse-etlHappy to answer questions about the design or trade-offs.

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Data infrastructure
Audience
Developer
AI stance
Not AI
Project type
Commercial product
Normalized one-liner
event transformations for clickhouse
Manually corrected
False

Could you build this?

No Sustaining real-time stream processing and transformations at 500,000+ events per second into ClickHouse requires deep distributed systems, memory management, and network I/O optimization.

What it would actually take: Requires an engine built in Rust, C++, or Go utilizing zero-allocation deserialization (e.g., SIMD-JSON), lock-free ring buffers, backpressure control, and batching mechanisms tailored to ClickHouse block storage. Requires deep performance engineering across Linux kernel networking (eBPF/DPDK), high-throughput message brokers (Kafka/Redpanda), and rigorous memory profiling to prevent garbage collection pauses and ingestion drops.

Discussion

4 comments analyzed.

Competitors mentioned: Flink (with ClickHouse connector), Flink JDBC/official connectors

Concerns raised: Flink's JVM heap management complexity, Operational overhead of Flink (jobs, state backends, checkpointing), Generic sinks not optimized for ClickHouse (batching, small inserts), Flink cluster management complexity (TaskManagers/JobManagers)

Competitors

Other products that read as similar to this one — 31 launches clear the similarity bar, closest 8 shown.

Attention rank: #14 of 32 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 143 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a data infrastructure tool for Media & entertainment yet.