Humor Arena
Which frontier model is funniest?
Details
- External ID
- 49190622
- Source
- HN
- Company
- —
- Product
- Humor Arena
- Website domain
- laugh.so
- Launched
- Aug. 5, 2026
- Cohort
- —
- Upvotes
- 10
- Upvotes percentile
- 0.5772849462365591
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:32 a.m.
- Updated at
- Sept. 10, 2026, 5:32 a.m.
Description
What if you could measure humor?Well we've trained a model on our own dataset of ~50k human ratings to detect what jokes people find funniest. We know it's part objective, part subjective component. Subjective is out of our depth for now hahaThe main results: Fable 5 is funniest - beating the average model 67% of the time, with GPT 4o last at 17%.Other findings: - The models never refused to try, even with dark prompts - Thinking longer has a slight benefit - Absurdness correlates negatively with joke qualitySome methodology notes: - We benchmarked our model against the human majority and it agreed 72% of the time in a blind sample test. - We had 51 US adults rate the jokes, each blind to the models, with joke order randomized, and quality checked for attention and speed. - To rate some yourself visit https://pair.laugh.soThe full benchmark here:https://laugh.so/benchmarkAm taking requests if there's more research you want to see! Cheers
Enrichment
- Theme
- interactive quiz and trivia games
- Vertical
- Horizontal
- Function
- —
- Audience
- B2C
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- compare humor capabilities across ai models
- Manually corrected
- False
Could you build this?
Partial While the web benchmark dashboard is standard, the core value relies on an internal dataset of 50k human humor ratings and an extensively tuned evaluation model.
What it would actually take: The architecture involves an LLM orchestration layer running prompts through dozens of target models, followed by a fine-tuned reward/judge model scoring pairwise humor matchups. Building this requires acquiring proprietary datasets of human humor preference judgments, calibrating Bradley-Terry or Elo rating algorithms, and validating judge correlations against human panels.
Discussion
7 comments analyzed.
Competitors mentioned: OpenCV, Mistral
Concerns raised: Humor assessment is highly subjective and relative, Most LLM responses aren't actually funny, 10x inference cost increase, Difficulty measuring laughter/genuine humor response
Feature requests: Multi-language benchmark dimension, Measure which LLMs succeed at unintentional humor, Test non-US built models performance
Competitors
Other products that read as similar to this one — 42 launches clear the similarity bar, closest 8 shown.
Attention rank: #18 of 43 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 265 days after the earliest competitor.
- Selbstbild · hn · 2026-07-05 · 6 upvotes · similarity 0.47
- Funny Guys · ph · 2026-09-18 · 2 upvotes · similarity 0.44
- Echo · hn · 2026-07-23 · 484 upvotes · similarity 0.42
- ha · ph · 2026-09-15 · 5 upvotes · similarity 0.40
- Antics: Drop-in multiplayer for your AI-built games · hn · 2026-07-29 · 5 upvotes · similarity 0.39
- A Dad Joke Website · hn · 2026-04-05 · 8 upvotes · similarity 0.39
- GIF&TAKE · ph · 2026-09-20 · 1 upvotes · similarity 0.39
- Knock Knock Joke Generator · ph · 2026-09-14 · 2 upvotes · similarity 0.39
Other launches for this product
- No other launches for this product.