I indexed 8,643 BSides talks across 227 chapters and 6 continents
Details
- External ID
- 48015655
- Source
- HN
- Company
- —
- Product
- I indexed 8,643 BSides talks across 227 chapters and 6 continents
- Website domain
- allbsides.com
- Launched
- May 4, 2026
- Cohort
- —
- Upvotes
- 25
- Upvotes percentile
- 0.7705977382875606
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hi HN,I'm Roland, and for the past few weeks, I've been building AllBSides — a directory of every BSides conference talk uploaded to YouTube. As of today, 8,643 talks from 5,927 speakers across 227 chapters in 68 countries. Combined runtime is 280 days. The transcripts come to about 60 million words.The archive came together in stages:1. Manually map every BSides chapter's YouTube channel 2. Pull every video and transcript from Supabase 3. Run each transcript through Haiku for tag extraction (tools, topics, difficulty, team, talk style, research method, and much more) 4. Run results through Sonnet for categorization and dedup 5. Final pass goes through Opus for verification 6. Do a manual verification - at one time, the pipeline showed over 16k AI suggestions for manual verification. Today, most are resolved.Total LLM cost so far: about €200. The whole pipeline is rebuildable from scratch.Each talk gets its own page with embedded video, full transcript, speakers, tags, and "related talks." Each tool/framework/protocol/standard mentioned across the corpus gets its own page (3,968 distinct technologies tracked).Some interesting facts I gathered while building it:-(A) The site is currently 94% bot traffic. Of that, about 80,000 hits/month are AI training crawlers (ClaudeBot, GPTBot, meta-externalagent). Within 7 days of the talks archive going live, all major AI labs had ingested the entire corpus. The discovery cascade was startling to watch in real time.-(B) The taxonomy work was the hardest part. Distinguishing "tools" from "frameworks" from "protocols" from "concepts" sounds easy until you have 5,000 ambiguous extracted entities. The 3-tier LLM pipeline helped a lot — Haiku alone was too noisy, Opus alone was too expensive.-(C) Top tools mentioned: Wireshark (343), PowerShell (342), Metasploit (332), Burp Suite (322), GitHub (296), VirusTotal (273), Docker (253), Splunk (251), Nmap (247), MITRE ATT&CK (237). The list reflects what BSides talks actually discuss, not what vendors curate.-(D) May is the peak BSides month — 29 events, 17% of all events with dates.-(E) The top 1% of talks (86 videos by view count) account for 51% of all viewership. The other 99% are deeply niche, often the only video record of a specific technique.The stack is intentionally lean: Go, SQLite, vanilla JavaScript, BunnyCDN. Static rendering at build time. No frameworks, no client-side state. The site costs about €50/month to run.The data behind this post and much more can be found in the site footer, under the link "stats".Happy to answer questions about the data pipeline, the taxonomy decisions, or what the AI crawler patterns looked like as the archive went live. Feedback on what to build next is genuinely welcome — I'm a solo dev figuring this out as I go.— Roland (parkado)
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Security
- Function
- Search & retrieval
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- index of bsides security talks
- Manually corrected
- False
Could you build this?
Yes This is a web directory aggregating YouTube metadata, speech-to-text transcripts, and full-text search, which can be rapidly built with off-the-shelf APIs and web frameworks.
Discussion
8 comments analyzed.
Competitors mentioned: RSA Conference, BlackHat
Concerns raised: Text readability/hard to read, Training data for LLMs/AI models, Security techniques becoming widely available in LLM training sets
Feature requests: Add BSides Perth to upcoming events list
Competitors
Other products that read as similar to this one — 63 launches clear the similarity bar, closest 8 shown.
Attention rank: #18 of 64 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 174 days after the earliest competitor.
- A searchable, timestamped index of 1,124 AI Engineer talks · hn · 2026-09-03 · 10 upvotes · similarity 0.54
- I mapped 8.5M research papers into an interactive atlas · hn · 2026-07-09 · 85 upvotes · similarity 0.42
- Visualizing How Books Reference Each Other Across 3k Years · hn · 2026-02-11 · 5 upvotes · similarity 0.40
- I used Claude Code to discover connections between 100 books · hn · 2026-01-10 · 524 upvotes · similarity 0.40
- Librarian · hn · 2026-02-26 · 8 upvotes · similarity 0.39
- I built a tiny LLM to demystify how language models work · hn · 2026-04-06 · 915 upvotes · similarity 0.39
- Was tired of drowning in HN comments, so I built an AI Chief of Staff · hn · 2026-01-26 · 5 upvotes · similarity 0.38
- TuringDB · hn · 2026-01-28 · 7 upvotes · similarity 0.38
Other launches for this product
- No other launches for this product.