Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Real-time avatars that change emotions as you talk

Details

External ID
48907348
Source
HN
Company
—
Product
Real-time avatars that change emotions as you talk
Website domain
anam.ai
Launched
July 14, 2026
Cohort
—
Upvotes
12
Upvotes percentile
0.6039426523297491
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hey HN, we're Ben and Caoimhe, cofounders of Anam. We build interactive avatars and just shipped our latest model, cara-4. This is the first avatar model which can naturally shift emotions and expressions during the conversation; a feature we’re calling “Director Notes”. The way it works is relatively simple: we have an LLM provide cues such as [laughter], [sad], [warm] interleaved with the speech, which we then condition our animation model with. To test it out, we commissioned a blind study from Mabyduck.com with 200 participants, 1,600 rated live interactions across six criteria. Cara-4 ranked first overall and was preferred head-to-head over each competitor on overall experience, lip-sync, visual quality and “naturalness”. On latency, measured across a full week of live traffic, end of user-speech to first video frame is ~1.2s median. The avatar model's own share is just ~100ms; most of the rest is waiting on STT, LLM, TTS or various forms of buffering (an unsung latency killer). How the model works: cara-4 has a two-stage design, a diffusion transformer turns audio+text into motion embeddings (head pose, gaze, lip shape, expression), and a rendering model applies those to a reference image, so new faces works without finetuning. Why faces at all: they carry emotional signal that text and voice don't, and they're a more accessible medium. Anam started in part from Ben watching his gran struggle with her iPad and thinking there should be a face she could just talk to. If you’d like to test it for free go to anam.ai

Enrichment

Theme
audio and signal processing tools
Vertical
Horizontal
Function
Hardware & robotics
Audience
B2C
AI stance
AI feature
Project type
Commercial product
Normalized one-liner
avatars that show emotions during conversation
Manually corrected
False

Could you build this?

No Real-time streaming conversational avatars with dynamic emotional shifts require proprietary diffusion/NeRF video synthesis models and ultra-low-latency real-time video generation infrastructure.

What it would actually take: This requires a proprietary generative neural rendering pipeline (custom 3D morphable models or diffusion-based audio-to-video synthesis) executing sub-200ms frame rendering on dedicated GPU clusters (e.g. H100s via TensorRT). WebRTC pipelines must synchronize incoming streaming audio, transcription, LLM orchestration, and frame generation with dynamic emotional conditioning latents. It requires a dedicated team of deep learning research scientists, computer vision experts, and real-time streaming engineers.

Discussion

10 comments analyzed.

Concerns raised: Latency dominated by STT, LLM, TTS components rather than avatar rendering, User connection adds significant delay to response time

Feature requests: Public evaluation report with experiment structure details

Competitors

Other products that read as similar to this one — 155 launches clear the similarity bar, closest 8 shown.

Attention rank: #66 of 156 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 257 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a hardware & robotics tool for Fintech yet.