Project AELLA
Open LLMs for structuring 100M research papers
Details
- External ID
- 45891092
- Source
- HN
- Company
- —
- Product
- Project AELLA
- Website domain
- inference.net
- Launched
- Nov. 11, 2025
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.27074235807860264
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
We're releasing Project AELLA - an open-science initiative to make scientific knowledge more accessible through AI-generated structured summaries of research papers.Blog: https://inference.net/blog/project-aellaVisualizer: https://aella.inference.netModels: https://huggingface.co/inference-net/Aella-Qwen3-14B, https://huggingface.co/inference-net/Aella-Nemotron-12BHighlights: - Released 100K research paper summaries in standardized JSON format with interactive visualization.- Fine-tuned open models (Qwen 3 14B & Nemotron 12B) that match GPT-5/Claude 4.5 performance at 98% lower cost (~$100K vs $5M to process 100M papers)- Built on distributed "idle compute" infrastructure - think SETI@Home for LLM workloadsGoal: Process ~100M papers total, then link to OpenAlex metadata and convert to copyright-respecting "Knowledge Units"The models are open, evaluation framework is transparent, and we're making the summaries publicly available. This builds on Project Alexandria's legal/technical foundation for extracting factual knowledge while respecting copyright.Technical deep-dive in the post covers our training pipeline, dual evaluation methods (LLM-as-judge + QA dataset), and economic comparison showing 50x cost reduction vs closed models.Happy to answer questions about the training approach, evaluation methodology, or infrastructure!
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Education
- Function
- Data infrastructure
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- open llms for structuring research papers
- Manually corrected
- False
Could you build this?
No Fine-tuning models and generating structured scientific extraction across 100 million research papers demands massive compute infrastructure and specialized ML data engineering.
What it would actually take: Requires a distributed data ingestion pipeline (e.g., Apache Spark or Ray) parsing petabytes of PDFs, a GPU cluster running distributed LLM fine-tuning on academic literature, and high-throughput inference infrastructure (such as vLLM) capable of processing 100M documents into structured knowledge graphs.
Discussion
2 comments analyzed.
Concerns raised: Naming conflict requiring product rename
Competitors
Other products that read as similar to this one — 128 launches clear the similarity bar, closest 8 shown.
Attention rank: #102 of 129 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 11 days after the earliest competitor.
- Run 500B+ Parameter LLMs Locally on a Mac Mini · hn · 2026-03-09 · 17 upvotes · similarity 0.44
- Data Engineering Book · hn · 2026-02-13 · 251 upvotes · similarity 0.43
- I mapped 8.5M research papers into an interactive atlas · hn · 2026-07-09 · 85 upvotes · similarity 0.43
- OpenFable · hn · 2026-04-08 · 5 upvotes · similarity 0.42
- litsurvey · github · 2026-09-11 · 7 upvotes · similarity 0.41
- Open Benchmarks Grants– a $3M commitment to close the AI eval gap · hn · 2026-02-11 · 6 upvotes · similarity 0.41
- Framework for building multi-agent equity research agents · hn · 2026-02-25 · 6 upvotes · similarity 0.41
- TabPFN-2.5 · hn · 2025-11-06 · 73 upvotes · similarity 0.40
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a data infrastructure tool for Media & entertainment yet.