I am building a map of people who lived in the Roman Empire
Details
- External ID
- 48481400
- Source
- HN
- Company
- —
- Product
- I am building a map of people who lived in the Roman Empire
- Website domain
- roman-names.com
- Launched
- June 10, 2026
- Cohort
- —
- Upvotes
- 210
- Upvotes percentile
- 0.9603825136612022
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Driving home from work one day, I wanted to know how many people we knew the names of who lived during the Roman era. Searching around, I found lists of Consuls and officials, but nothing that covered ordinary people or even most people like freedmen and slaves. So I ended up building a pipeline to process the more than 500k Latin inscriptions in the Epigraphic Database Clauss-Slaby https://edcs.hist.uzh.ch/en/ and extract the names of people (and attempt to cluster them, but this is a work in progress).There are databases where Classicists have done this manually for specific regions, Trismegistos https://www.trismegistos.org/ and Latin Inscriptions of the Roman Empire (LIRE) https://pure.au.dk/portal/en/publications/latin-inscriptions... are two major efforts I found. But there doesn't seem to be a project that did what I set out to do, although I have read in some places that it was believed to be possible.I am not a classicist or a web developer, but I have Claude and Gemini and I can sort of read basic Latin - so I set to work. I used LIRE and another database as ground truth and built a pipeline to extract and process the inscriptions to recover the names. The process I developed uses a high end LLM like Sonnet or Gemini Pro to supervise the extraction and tuning process on a regional basis until the obvious error rate is reasonable. For this, so far, reasonable to me means less than 1-2% in the smaller initial samples of 100-500 and no observed systemic issues. The different regions often need different prompts, so this basically became an exercise in letting the higher level AI tune the prompt for the lower level AI. The extraction when measured against LIRE produces an F1 score between 0.64 and 0.87, but take this with a grain of salt.Once I had done a few regions, I wanted to see the work, so I threw together a pretty crude website but as I am not a web developer, it was crude in how it accessed its data. It does look cool and I also added summarization, and machine translation to each entry. I wanted to eventually get feedback from an actual team of classicists and make the website work better, so I am rewriting it as we speak but it is broadly functional now with a few extra bugs but substantially improved performance compared to the old one. All entries link back to the proper sources, and the old web app linked to several additional sources where the data was present, but I haven't gotten that working again just yet on the new one. (The old web interface is still available at https://roman-names.com, but I will warn you it is clunky and not mobile friendly at all)Key findings so far:AI supervised AI extraction saved me time. I was manually tuning things for a while and then the runbook became an idea that I feed my instructions in and let the big AI go with sparse oversight from me.The extraction improved significantly (by about 10 F1 points) when I fed the model the raw text including the markers, vs a cleaned up version of the text.I just thought it was a cool little project and wanted to share. If you happen to work in any adjacent space and there is something I could do better etc let me know.
Enrichment
- Theme
- astrology, spirituality, and lifestyle tools
- Vertical
- Horizontal
- Function
- Analytics & BI
- Audience
- B2C
- AI stance
- —
- Project type
- Hobby / open-source project
- Normalized one-liner
- map of people who lived in the roman empire
- Manually corrected
- False
Could you build this?
Partial The Leaflet/MapLibre map UI is easily vibe-coded, but the historical data pipeline requires parsing, standardizing, and geocoding heterogeneous classical epigraphy databases (like EDCS or Clauss-Slaby).
What it would actually take: A complete system needs an ETL pipeline ingesting academic epigraphic datasets (Epigraphic Database Clauss-Slaby, Trismegistos, EDH) with custom parsers for Latin onomastics (praenomina, gentilitia, cognomina), entity deduplication, and gazetteer-based georeferencing (Pleiades). The frontend can be static GeoJSON on MapLibre GL, but domain-specific parsing of fragmentary ancient inscriptions is the difficult component.
Discussion
20 comments analyzed.
Competitors mentioned: Greek Lexicon of Personal Names (LGPN), EDCS (Epigraphic Database of Latin Inscriptions)
Concerns raised: Zoom-level-dependent clustering obscures local distribution of points, Imprecision in LLM-generated transcriptions and translations, AI hallucinations in name extraction and Latin phrase translation, Poor geolocation data for 1800s finds limits spatial accuracy, Unclear definition of 'status' field in data
Feature requests: Connect people via relationship links (wife of, etc.), Include EDCS source references for tracking original inscriptions, Compare LLM model performance (Gemini 3.5/3.1 vs current Flash-Lite)
Competitors
Other products that read as similar to this one — 42 launches clear the similarity bar, closest 8 shown.
Attention rank: #3 of 43 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 219 days after the earliest competitor.
- Help SourceLibrary.org Translate the Renaissance · hn · 2026-06-07 · 9 upvotes · similarity 0.46
- The Roman Industrial Revolution that could have been (Vol 2) · hn · 2026-03-06 · 43 upvotes · similarity 0.42
- Ex Situ · hn · 2026-07-21 · 51 upvotes · similarity 0.40
- A map of historical movies by narrative location and time period · hn · 2026-02-01 · 7 upvotes · similarity 0.40
- Zenòdot · hn · 2026-03-09 · 15 upvotes · similarity 0.38
- Visualizing How Books Reference Each Other Across 3k Years · hn · 2026-02-11 · 5 upvotes · similarity 0.38
- Large Scale Article Extract of Newspapers 1730s-1960s · hn · 2026-05-02 · 57 upvotes · similarity 0.37
- I built Chronoscope, because Google Maps won't let you visit 3400 BCE · hn · 2026-03-12 · 11 upvotes · similarity 0.36
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a analytics & bi tool for Legal yet.