Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Cancer diagnosis makes for an interesting RL environment for LLMs

Details

External ID
45902590
Source
HN
Company
—
Product
—
Website domain
—
Launched
Nov. 12, 2025
Cohort
—
Upvotes
46
Upvotes percentile
0.8198689956331878
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hey HN, this is David from Aluna (YC S24). We work with diagnostic labs to build datasets and evals for oncology tasks.I wanted to share a simple RL environment I built that gave frontier LLMs a set of tools that lets it zoom and pan across a digitized pathology slide to find the relevant regions to make a diagnosis. Here are some videos of the LLM performing diagnosis on a few slides:(https://www.youtube.com/watch?v=k7ixTWswT5c): traces of an LLM choosing different regions to view before making a diagnosis on a case of small-cell carcinoma of the lung(https://youtube.com/watch?v=0cMbqLnKkGU): traces of an LLM choosing different regions to view before making a diagnosis on a case of benign fibroadenoma of the breastWhy I built this:Pathology slides are the backbone of modern cancer diagnosis. Tissue from a biopsy is sliced, stained, and mounted on glass for a pathologist to examine abnormalities.Today, many of these slides are digitized into whole-slide images (WSIs)in TIF or SVS format and are several gigabytes in size.While there exists several pathology-focused AI models, I was curious to test whether frontier LLMs can perform well on pathology-based tasks. The main challenge is that WSIs are too large to fit into an LLM’s context window. The standard workaround, splitting them into thousands of smaller tiles, is inefficient for large frontier LLMs.Inspired by how pathologists zoom and pan under a microscope, I built a set of tools that let LLMs control magnification and coordinates, viewing small regions at a time and deciding where to look next.This ended up resulting in some interesting behaviors, and actually seemed to yield pretty good results with prompt engineering:- GPT 5: explored up to ~30 regions before deciding (concurred with an expert pathologist on 4 out of 6 cancer subtyping tasks and 3 out of 5 IHC scoring tasks)- Claude 4.5: Typically used 10–15 views but similar accuracy as GPT-5 (concurred with the pathologist on 3 out of 6 cancer subtyping tasks and 4 out of 5 IHC scoring tasks)- Smaller models (GPT 4o, Claude 3.5 Haiku): examined ~8 frames and were less accurate overall (1 out of 6 cancer subtytping tasks and 1 out of 5 IHC scoring tasks)Obviously, this was a small sample set, so we are working on creating a larger benchmark suite with more cases and types of tasks, but I thought this was cool that it even worked so I wanted to share with HN!

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Healthcare
Function
Dev tools
Audience
Developer
AI stance
AI feature
Project type
Hobby / open-source project
Normalized one-liner
cancer diagnosis reinforcement learning environment
Manually corrected
False

Could you build this?

Partial The LLM tool wrapper is vibe-codeable, but creating the whole-slide pathology environment requires handling gigapixel multi-resolution WSI files and specialized medical imaging libraries.

What it would actually take: The system requires a gym-like RL environment built on OpenSlide/PyTorch to serve pyramid-tiled gigapixel pathology images (SVS/NDPI formats), managing dynamic pan/zoom coordinate transforms, coupled with domain-specific pathology evaluation metrics to reward the LLM's diagnostic search trajectories.

Discussion

20 comments analyzed.

Competitors mentioned: Prov-GigaPath, CTransPath, Pathology-specific foundation models with ViT backbones, Specialized point-solution FDA-approved tools

Concerns raised: Unclear business model and go-to-market strategy, Lacks access to huge labeled whole-slide image datasets like competitors, Not accurate enough for clinical deployment yet, Unlimited range of possible diagnoses makes generalization difficult, 90% accuracy insufficient for FDA approval and clinical use

Feature requests: Fine-tuning on different modalities (IMC, H&E) for better generalization, Explore newer segmentation models for improved performance, Better evaluation methodology separating navigation from assessment capabilities, Implement video compression techniques for memory management optimization

Competitors

Other products that read as similar to this one — 30 launches clear the similarity bar, closest 8 shown.

Attention rank: #8 of 31 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 8 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a dev tools tool for Sales yet.