Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Humor Arena

Which frontier model is funniest?

Details

External ID
49190622
Source
HN
Company
—
Product
Humor Arena
Website domain
laugh.so
Launched
Aug. 5, 2026
Cohort
—
Upvotes
10
Upvotes percentile
0.5772849462365591
Tags
—
Fetched at
Sept. 10, 2026, 5:32 a.m.
Updated at
Sept. 10, 2026, 5:32 a.m.

Description

What if you could measure humor?Well we've trained a model on our own dataset of ~50k human ratings to detect what jokes people find funniest. We know it's part objective, part subjective component. Subjective is out of our depth for now hahaThe main results: Fable 5 is funniest - beating the average model 67% of the time, with GPT 4o last at 17%.Other findings: - The models never refused to try, even with dark prompts - Thinking longer has a slight benefit - Absurdness correlates negatively with joke qualitySome methodology notes: - We benchmarked our model against the human majority and it agreed 72% of the time in a blind sample test. - We had 51 US adults rate the jokes, each blind to the models, with joke order randomized, and quality checked for attention and speed. - To rate some yourself visit https://pair.laugh.soThe full benchmark here:https://laugh.so/benchmarkAm taking requests if there's more research you want to see! Cheers

Enrichment

Theme
interactive quiz and trivia games
Vertical
Horizontal
Function
—
Audience
B2C
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
compare humor capabilities across ai models
Manually corrected
False

Could you build this?

Partial While the web benchmark dashboard is standard, the core value relies on an internal dataset of 50k human humor ratings and an extensively tuned evaluation model.

What it would actually take: The architecture involves an LLM orchestration layer running prompts through dozens of target models, followed by a fine-tuned reward/judge model scoring pairwise humor matchups. Building this requires acquiring proprietary datasets of human humor preference judgments, calibrating Bradley-Terry or Elo rating algorithms, and validating judge correlations against human panels.

Discussion

7 comments analyzed.

Competitors mentioned: OpenCV, Mistral

Concerns raised: Humor assessment is highly subjective and relative, Most LLM responses aren't actually funny, 10x inference cost increase, Difficulty measuring laughter/genuine humor response

Feature requests: Multi-language benchmark dimension, Measure which LLMs succeed at unintentional humor, Test non-US built models performance

Competitors

Other products that read as similar to this one — 42 launches clear the similarity bar, closest 8 shown.

Attention rank: #18 of 43 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 265 days after the earliest competitor.

Other launches for this product