Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

I built "AI Wattpad" to eval LLMs on fiction

Details

External ID
46873742
Source
HN
Company
—
Product
I built "AI Wattpad" to eval LLMs on fiction
Website domain
narrator.sh
Launched
Feb. 3, 2026
Cohort
—
Upvotes
32
Upvotes percentile
0.7513477088948787
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

I've been a webfiction reader for years (too many hours on Royal Road), and I kept running into the same question: which LLMs actually write fiction that people want to keep reading? That's why I built Narrator (https://narrator.sh/llm-leaderboard) – a platform where LLMs generate serialized fiction and get ranked by real reader engagement.Turns out this is surprisingly hard to answer. Creative writing isn't a single capability – it's a pipeline: brainstorming → writing → memory. You need to generate interesting premises, execute them with good prose, and maintain consistency across a long narrative. Most benchmarks test these in isolation, but readers experience them as a whole.The current evaluation landscape is fragmented: Memory benchmarks like FictionLive's tests use MCQs to check if models remember plot details across long contexts. Useful, but memory is necessary for good fiction, not sufficient. A model can ace recall and still write boring stories.Author-side usage data from tools like Novelcrafter shows which models writers prefer as copilots. But that measures what's useful for human-AI collaboration, not what produces engaging standalone output. Authors and readers have different needs.LLM-as-a-judge is the most common approach for prose quality, but it's notoriously unreliable for creative work. Models have systematic biases (favoring verbose prose, certain structures), and "good writing" is genuinely subjective in ways that "correct code" isn't.What's missing is a reader-side quantitative benchmark – something that measures whether real humans actually enjoy reading what these models produce. That's the gap Narrator fills: views, time spent reading, ratings, bookmarks, comments, return visits. Think of it as an "AI Wattpad" where the models are the authors.I shared an early DSPy-based version here 5 months ago (https://news.ycombinator.com/item?id=44903265). The big lesson: one-shot generation doesn't work for long-form fiction. Models lose plot threads, forget characters, and quality degrades across chapters.The rewrite: from one-shot to a persistent agent loopThe current version runs each model through a writing harness that maintains state across chapters. Before generating, the agent reviews structured context: character sheets, plot outlines, unresolved threads, world-building notes. After generating, it updates these artifacts for the next chapter. Essentially each model gets a "writer's notebook" that persists across the whole story.This made a measurable difference – models that struggled with consistency in the one-shot version improved significantly with access to their own notes.Granular filtering instead of a single score:We classify stories upfront by language, genre, tags, and content rating. Instead of one "creative writing" leaderboard, we can drill into specifics: which model writes the best Spanish Comedy? Which handles LitRPG stories with Male Leads the best? Which does well with romance versus horror?The answers aren't always what you'd expect from general benchmarks. Some models that rank mid-tier overall dominate specific niches.A few features I'm proud of:Story forking lets readers branch stories CYOA-style – if you don't like where the plot went, fork it and see how the same model handles the divergence. Creates natural A/B comparisons.Visual LitRPG was a personal itch to scratch. Instead of walls of [STR: 15 → 16] text, stats and skill trees render as actual UI elements. Example: https://narrator.sh/novel/beware-the-starter-pet/chapter/1What I'm looking for:More readers to build out the engagement data. Also curious if anyone else working on long-form LLM generation has found better patterns for maintaining consistency across chapters – the agent harness approach works but I'm sure there are improvements.

Enrichment

Theme
ai storytelling and children's education
Vertical
Horizontal
Function
Observability & eval
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
llm evaluation on fiction
Manually corrected
False

Could you build this?

Yes It is a web application that calls LLM APIs to generate chapters, displays webfiction, and collects user engagement metrics to populate a leaderboard.

Discussion

20 comments analyzed.

Competitors mentioned: Claude (AI writing alternative), LMArena (similar benchmarking platform), Design Arena (live benchmark comparison), Wattpad (fiction platform comparison)

Concerns raised: Generated stories are incoherent and unreadable, AI writing lacks emotional believability and consistency, Output is mediocre, generic, and poorly written, Unclear how usable preference data can be collected from low-quality outputs, No differentiation from reading someone else's AI-generated fiction vs. personal AI writing

Feature requests: Personalized fiction generation based on individual user tastes, Human-AI collaborative system for plot brainstorming, Improve consistency across chapters and paragraphs, Story remixing/forking functionality (already mentioned as planned)

Competitors

Other products that read as similar to this one — 140 launches clear the similarity bar, closest 8 shown.

Attention rank: #34 of 141 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 94 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a observability & eval tool for Media & entertainment yet.