Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

HelixDB

A graph database built on object storage

Details

External ID
48478148
Source
HN
Company
—
Product
HelixDB
Website domain
github.com
Launched
June 10, 2026
Cohort
—
Upvotes
159
Upvotes percentile
0.9467213114754098
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hey HN, it’s been just over a year since we launched HelixDB (https://news.ycombinator.com/item?id=43975423), a project a friend and I started in college. It’s an OLTP graph database built on object-storage, with native vector search and full-text search (FTS).Why graph, vector and FTS? Graph databases provide a natural cognitive model for data, vectors allow for a semantic understanding of the entities and relationships in the graph, and FTS provides more specific filtering. Many AI-driven applications attempt to combine all of these functionalities by stitching together multiple disconnected systems, but even then there’s no native way to perform joins or queries that span all systems. You still need to handle this logic at the application level.Helix started as a graph DB, but we moved to a hybrid graph/vector approach after attempting to build an AI memory system, which led us down the GraphRAG and HybridRAG rabbit hole, where we would need separate graph and vector databases.We knew scalability would be a challenge at each stage of our product's development, however our initial focus this past year was to prove out the product through local deployments and was only meant to be run on a single node. Scaling graph DBs remained a difficult and expensive problem we’d have to solve later. Some common ways other graph DBs solve scaling is by duplicating entire datasets across distributed machines (extremely expensive per node), or by sharding the data.Sharding databases is effective and affordable, however, graph data doesn’t have explicit partitions like relational databases do. For example, sharding a relational DB involves splitting up tables. When it comes to graph DBs, the edges can span across any of the partitions, and hopping across multiple machines when traversing nodes is ineffective and computationally expensive.Replicating graph DBs for high availability and better throughput drastically increases the operational cost of the db and still has a limit of how big you can vertically scale. The workload that we’re used for requires storing a huge amount of data for agents, where only a subset of that data is ever needed at any one time. So rather than having the whole thing in memory, we can store it all in object-storage and get the bits we need when they’re needed.Agents benefit from better context, which is achieved from more and better data (more relationships etc). By using S3 as the persistence/data layer there is no limit to how big the graph can be or how many relationships you can have, and we can scale to serve throughput and requests by horizontally spinning up nodes and caching relevant subsets of the graph on each node. This way, you get extremely low latency for “hot” data and a p99 of ~100ms for writes and ~50ms for reads from cold storage (S3). Plus you get the benefit of dirt cheap storage.Workloads that HelixDB is currently supporting: - Huge amounts of data (TBs) from which the agents need to search and traverse over - Offering affordable graph storage for companies where cost of graph data is a bottleneck - Consolidating multiple databases, enabling AI agents to have autonomy over companies, helping them become more autonomous. - AI memory - Company brainsWe’re currently working on our own generalised AI memory layer which will use HelixDB under the hood and be completely open-source. Also, we’re finishing up on pre-filtering for vector search which will allow you to pre-filter based on relationships in the graph, metadata, and sub-graphs. And lastly, GA cloud will be available in the coming weeks.If you want to run Helix locally (either on-disk or in-memory), you can find more info on our github (https://github.com/HelixDB/helix-db) or via our docs (https://docs.helix-db.com/database/local-development). If you’re interested in getting started with our distributed cloud, please email us [email protected] thanks! Comments and feedback welcome!

Enrichment

Theme
database infrastructure and developer tools
Vertical
Horizontal
Function
Data infrastructure
Audience
Developer
AI stance
Not AI
Project type
Commercial product
Normalized one-liner
graph database on object storage
Manually corrected
False

Could you build this?

No Building an ACID OLTP graph database natively on top of object storage with integrated vector search and inverted index FTS requires world-class database systems and storage engine engineering.

What it would actually take: Requires implementing a storage engine in Rust/C++ optimized for high-latency S3-compatible APIs using LSM-tree or partitioned B-tree variants, distributed concurrency control (MVCC), an execution engine for graph traversal, custom SIMD-accelerated vector indexing (like HNSW), and a full-text search engine. Writing distributed storage systems that ensure consistency and performant OLTP latencies over object storage demands seasoned systems programmers.

Discussion

20 comments analyzed.

Competitors mentioned: PostgreSQL, PuppyGraph, Other cloud object stores with S3 API

Concerns raised: High latency for deep multi-hop queries in cold storage (~50ms per hop), Cloud pricing too expensive for experimentation ($600/mo minimum), Not open source / source code access limited, Developer hesitation adopting new data stores

Feature requests: Self-hosted free or low-cost option, Open source code availability, Faster query performance for deep multi-hop traversals

Competitors

Other products that read as similar to this one — 159 launches clear the similarity bar, closest 8 shown.

Attention rank: #11 of 160 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 224 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a data infrastructure tool for Media & entertainment yet.