Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

We fingerprinted 178 AI models' writing styles and similarity clusters

Details

External ID
47690415
Source
HN
Company
—
Product
We fingerprinted 178 AI models' writing styles and similarity clusters
Website domain
rival.tips
Launched
April 8, 2026
Cohort
—
Upvotes
78
Upvotes percentile
0.877892030848329
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

We have a dataset of 3,095 standardized AI responses across 43 prompts. From each response, we extract a 32-dimension stylometric fingerprint (lexical richness, sentence structure, punctuation habits, formatting patterns, discourse markers).Some findings:- 9 clone clusters (>90% cosine similarity on z-normalized feature vectors) - Mistral Large 2 and Large 3 2512 score 84.8% on a composite metric combining 5 independent signals - Gemini 2.5 Flash Lite writes 78% like Claude 3 Opus. Costs 185x less - Meta has the strongest provider "house style" (37.5x distinctiveness ratio) - "Satirical fake news" is the prompt that causes the most writing convergence across all models - "Count letters" causes the most divergenceThe composite clone score combines: prompt-controlled head-to-head similarity, per-feature Pearson correlation across challenges, response length correlation, cross-prompt consistency, and aggregate cosine similarity.Tech: stylometric extraction in Node.js, z-score normalization, cosine similarity for aggregate, Pearson correlation for per-feature tracking. Analysis script is ~1400 lines.

Enrichment

Theme
AI text humanizers and detectors
Vertical
Security
Function
Observability & eval
Audience
Developer
AI stance
AI feature
Project type
Hobby / open-source project
Normalized one-liner
ai model writing style fingerprinting analysis
Manually corrected
False

Could you build this?

Partial Extracting stylometric features from text via standard NLP libraries is straightforward, but assembling, cleaning, and normalizing a benchmark dataset across 178 models requires substantial automated evaluation pipelines and API compute.

What it would actually take: The implementation requires automated prompt evaluation harnesses querying dozens of model endpoints via standard APIs (OpenAI, Anthropic, open-source model providers via vLLM), followed by a feature extraction pipeline computing 32 linguistic and statistical metrics (lexical diversity via TTR/Yule's K, POS distribution via spaCy, sentence entropy, and formatting tokens). Dimensionality reduction and clustering (UMAP, HDBSCAN, or hierarchical clustering) must then be applied and validated across thousands of samples.

Discussion

20 comments analyzed.

Competitors mentioned: models.dev, arena.ai, HuggingFace

Concerns raised: Methodology lacks linguistic theory and uses arbitrary metrics/thresholds, Claims (75% similarity = 'writes the same') unsubstantiated without prompts/responses shown, Accessibility issues: muted colors on dark background, poor contrast ratios, Suspicious claim that Opus and Gemini Flash share 99% style similarity, No explanation of how 32 dimensions were selected or if prompts were optimized

Competitors

Other products that read as similar to this one — 227 launches clear the similarity bar, closest 8 shown.

Attention rank: #26 of 228 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 159 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a observability & eval tool for Media & entertainment yet.