Postgres extension for BM25 relevance-ranked full-text search
Details
- External ID
- 47589856
- Source
- HN
- Company
- —
- Product
- Postgres extension for BM25 relevance-ranked full-text search
- Website domain
- github.com
- Launched
- March 31, 2026
- Cohort
- —
- Upvotes
- 203
- Upvotes percentile
- 0.9630996309963099
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Last summer we faced a conundrum at my company, Tiger Data, a Postgres cloud vendor whose main business is in timeseries data. We were trying to grow our business towards emerging AI-centric workloads and wanted to provide a state-of-the-art hybrid search stack in Postgres. We'd already built pgvectorscale in house with the goal of scaling semantic search beyond pgvector's main memory limitations. We just needed a scalable ranked keyword search solution too.The problem: core Postgres doesn't provide this; the leading Postgres BM25 extension, ParadeDB, is guarded behind AGPL; developing our own extension appeared daunting. We'd need a small team of sharp engineers and 6-12 months, I figured. And we'd probably still fall short of the performance of a mature system like Parade/Tantivy.Or would we? I'd be experimenting long enough with AI-boosted development at that point to realize that with the latest tools (Claude Code + Opus) and an experienced hand (I've been working in database systems internals for 25 years now), the old time estimates pretty much go out the window.I told our CTO I thought I could solo the project in one quarter. This raised some eyebrows.It did take a little more time than that (two quarters), and we got some real help from the community (amazing!) after open-sourcing the pre-release. But I'm thrilled/exhausted today to share that pg_textsearch v1.0 is freely available via open source (Postgres license), on Tiger Data cloud, and hopefully soon, a hyperscalar near you:https://github.com/timescale/pg_textsearchIn the blog post accompanying the release, I overview the architecture and present benchmark results using MS-MARCO. To my surprise, we were not only able to meet Parade/Tantivy's query performance, but exceed it substantially, measuring a 4.7x advantage on query throughput at scale:https://www.tigerdata.com/blog/pg-textsearch-bm25-full-text-...It's exciting (and, to be honest, a little unnerving) to see a field I've spent so much time toiling in change so quickly in ways that enable us to be more ambitious in our technical objectives. Technical moats are moats no longer.The benchmark scripts and methodology are available in the github repo. Happy to answer any questions in the thread.Thanks,TJ ([email protected])
Enrichment
- Theme
- database infrastructure and developer tools
- Vertical
- Horizontal
- Function
- Search & retrieval
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- postgres extension for bm25 full-text search
- Manually corrected
- False
Could you build this?
No Authoring a native C/Rust PostgreSQL extension that implements custom index access methods for BM25 ranking requires deep PostgreSQL internals and information retrieval systems knowledge.
What it would actually take: The project requires writing a Postgres C extension using custom Index Access Methods (TAM/AM) or specialized inverted indexes (GIN/block-level inverted lists) adhering to Postgres storage and WAL recovery conventions. It requires deep knowledge of database internals, concurrency control (latches/locks), and optimized BM25 scoring algorithms across billions of terms.
Discussion
20 comments analyzed.
Competitors mentioned: pgvectorscale (Rust-based alternative), Tantivy (search engine), ParadeDB, Elasticsearch, Oban (job processing)
Concerns raised: Only supports PostgreSQL 17, not PG18, Timestamp with time zone doesn't actually store timezone information, Case-insensitive search may be better solved with citext, Separate search index preferable to single source of truth for large-scale use, ETL pipeline and data reindexing strategies need careful architecture
Feature requests: Support for Azure Flexible Server, Prefix search support for fields like lastname, Reciprocal rank fusion (RRF) for hybrid lexical and vector search, Better integration with fuzzy search and phrase matching
Competitors
Other products that read as similar to this one — 54 launches clear the similarity bar, closest 8 shown.
Attention rank: #3 of 55 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 153 days after the earliest competitor.
- Semantic search over Hacker News, built on pgvector · hn · 2026-02-22 · 5 upvotes · similarity 0.50
- lead · github · 2026-09-17 · 144 upvotes · similarity 0.45
- Widen · hn · 2026-07-31 · 7 upvotes · similarity 0.42
- SereneDB · ph · 2026-09-22 · 197 upvotes · similarity 0.42
- ErwinDB, a TUI to view 7k Stack Overflow answers by Postgres expert · hn · 2026-02-03 · 5 upvotes · similarity 0.40
- CS · hn · 2026-02-22 · 15 upvotes · similarity 0.40
- Pgrust, Postgres in Rust (passing 100% of Postgres regression tests) · hn · 2026-06-25 · 6 upvotes · similarity 0.38
- pg-jev · github · 2026-09-17 · 274 upvotes · similarity 0.38
Other launches for this product
- No other launches for this product.