500k+ events/sec transformations for ClickHouse ingestion
Details
- External ID
- 47693407
- Source
- HN
- Company
- —
- Product
- 500k+ events/sec transformations for ClickHouse ingestion
- Website domain
- github.com
- Launched
- April 8, 2026
- Cohort
- —
- Upvotes
- 13
- Upvotes percentile
- 0.6658097686375322
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hi HN! We are Ashish and Armend, founders of GlassFlow.Over the last year, we worked with teams running high-throughput pipelines into self-hosted ClickHouse. Mostly for observability and real-time analytics.A question that came repeatedly was: What happens when throughput grows?Usually, things work fine at 10k events/sec, but we started seeing backpressure and errors at >100k.When the throughput per pipeline stops scaling, then adding more CPU/memory doesn’t help because often parts of the pipeline are not parallelized or are bottlenecked by state handling.At this point, engineers usually scale by adding more pipeline instances.That works but comes with some trade-offs: - You have to split the workload (e.g., multiple pipelines reading from the same source) - Transformation logic gets duplicated across pipelines - Stateful logic becomes harder to manage and keep consistent - Debugging and changes get more difficult because the data flow is fragmentedAnother challenge arises when working with high-cardinality keys like user IDs, session IDs, or request IDs, and when you need to handle longer time windows (24h or more). The state grows quickly and many systems rely on in-memory state, which makes it expensive and harder to recover from failures.We wanted to solve this problem and rebuild our approach at GlassFlow.Instead of scaling by adding more pipelines, we scale within a single pipeline by using replicas. Each replica consumes, processes, and writes independently, and the workload is distributed across them.In the benchmarks we’re sharing, this scales to 500k+ events/sec while still running stateful transformations and writing into ClickHouse.A few things we think are interesting: - Scaling is close to linear as you add replicas - Works with stateful transformations (not just stateless ingestion) - State is backed by a file-based KV store instead of relying purely on memory - The ClickHouse sink is optimized for batching to avoid small inserts - The product is built with GoFull write-up + benchmarks: https://www.glassflow.dev/blog/glassflow-now-scales-to-500k-...Repo: https://github.com/glassflow/clickhouse-etlHappy to answer questions about the design or trade-offs.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Data infrastructure
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- event transformations for clickhouse
- Manually corrected
- False
Could you build this?
No Sustaining real-time stream processing and transformations at 500,000+ events per second into ClickHouse requires deep distributed systems, memory management, and network I/O optimization.
What it would actually take: Requires an engine built in Rust, C++, or Go utilizing zero-allocation deserialization (e.g., SIMD-JSON), lock-free ring buffers, backpressure control, and batching mechanisms tailored to ClickHouse block storage. Requires deep performance engineering across Linux kernel networking (eBPF/DPDK), high-throughput message brokers (Kafka/Redpanda), and rigorous memory profiling to prevent garbage collection pauses and ingestion drops.
Discussion
4 comments analyzed.
Competitors mentioned: Flink (with ClickHouse connector), Flink JDBC/official connectors
Concerns raised: Flink's JVM heap management complexity, Operational overhead of Flink (jobs, state backends, checkpointing), Generic sinks not optimized for ClickHouse (batching, small inserts), Flink cluster management complexity (TaskManagers/JobManagers)
Competitors
Other products that read as similar to this one — 31 launches clear the similarity bar, closest 8 shown.
Attention rank: #14 of 32 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 143 days after the earliest competitor.
- ObsessionDB · hn · 2026-01-23 · 12 upvotes · similarity 0.48
- SensorFlow · ph · 2026-09-27 · 1 upvotes · similarity 0.44
- Ingestlayer · hn · 2026-06-22 · 9 upvotes · similarity 0.40
- WaveHouse · hn · 2026-08-20 · 20 upvotes · similarity 0.38
- Managed Postgres with native ClickHouse integration · hn · 2026-01-22 · 45 upvotes · similarity 0.36
- The TypeScript Semantic Layer for ClickHouse · hn · 2026-06-27 · 8 upvotes · similarity 0.36
- ghost · github · 2026-09-19 · 18 upvotes · similarity 0.35
- TabPFN Scaling Mode · hn · 2025-12-03 · 5 upvotes · similarity 0.35
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a data infrastructure tool for Media & entertainment yet.