Public Apache Iceberg datasets via a REST catalog
Details
- External ID
- 46636891
- Source
- HN
- Company
- —
- Product
- Public Apache Iceberg datasets via a REST catalog
- Website domain
- googleblog.com
- Launched
- Jan. 15, 2026
- Cohort
- —
- Upvotes
- 13
- Upvotes percentile
- 0.5790513833992095
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hi HN,I’m one of the creators of this project.We noticed that while many developers want to experiment with Apache Iceberg the "entry cost" is often high. You usually have to set up your own storage buckets, configure a catalog (like Hive), and ingest data before you can even run a single SELECT statement.We wanted to lower that barrier. We’ve hosted a production-grade Iceberg REST Catalog on BigLake with public datasets (starting with the NYC Taxi data) that anyone can query.You can point Spark, Trino, or Flink directly at the REST endpoint and start querying immediately.You do need a Google Cloud Project ID for authentication/quota, but the data access itself is free and public.I’d love to hear your thoughts. Are there specific datasets or Iceberg features you’d like to see added to the dataset?
Enrichment
- Theme
- database infrastructure and developer tools
- Vertical
- Horizontal
- Function
- Data infrastructure
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- apache iceberg dataset catalog for developers
- Manually corrected
- False
Could you build this?
No Providing a public Apache Iceberg REST catalog with hosted datasets requires managing substantial cloud infrastructure, high-throughput distributed object storage, and specialized data engineering at Google Cloud scale.
What it would actually take: A production version requires deploying an Apache Iceberg REST Catalog specification implementation backed by a high-availability metadata store (like PostgreSQL/FoundationDB) and multi-terabyte cloud object storage (GCS/S3). It involves building automated ETL pipelines for public datasets, handling egress costs, multi-tenant rate limiting, and integrating engines like BigLake/Trino. This requires deep distributed systems and cloud data engineering expertise.
Discussion
6 comments analyzed.
Competitors mentioned: Trino, Spark
Concerns raised: Python access to REST catalog and datasets unclear
Feature requests: Iceberg v3 spec support (Variant, Deletion Vector), Python example for querying datasets, Trino Docker config integration
Competitors
Other products that read as similar to this one — 60 launches clear the similarity bar, closest 8 shown.
Attention rank: #23 of 61 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 78 days after the earliest competitor.
- Iceberg-JS, a TypeScript Client for the Apache Iceberg REST Catalog · hn · 2025-12-09 · 5 upvotes · similarity 0.59
- IceGate · hn · 2026-04-13 · 15 upvotes · similarity 0.45
- Marmot · hn · 2025-12-02 · 103 upvotes · similarity 0.43
- datafusion-iceberg · github · 2026-09-09 · 8 upvotes · similarity 0.43
- 7x faster Iceberg ingestion, how we redesigned OLake's writer · hn · 2025-12-05 · 5 upvotes · similarity 0.41
- I built a local data lake for AI powered data engineering and analytics · hn · 2026-04-08 · 14 upvotes · similarity 0.40
- CloudClerk. We struggled with BigQuery finops, so we decided to fight · hn · 2026-01-23 · 5 upvotes · similarity 0.40
- StreamHouse · hn · 2026-02-25 · 10 upvotes · similarity 0.40
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a data infrastructure tool for Media & entertainment yet.