Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

InstinctFlash

High-Performance Serving Runtime for Robotics Models

Details

External ID
49802789
Source
HN
Company
—
Product
InstinctFlash
Website domain
github.com
Launched
Sept. 22, 2026
Cohort
—
Upvotes
27
Upvotes percentile
0.7814992025518341
Tags
—
Fetched at
Sept. 26, 2026, 10:53 p.m.
Updated at
Sept. 26, 2026, 10:53 p.m.

Description

Hey HN, Guanming here, cofounder of General Instinct. We just released InstinctFlash, a high-performance serving framework for robotics models. It’s licensed under AGPL-3.0.On Jetson Thor, we see speedups about 1.2x to 7.9x from runtime optimizations alone and up to 33.78x for LingBot-VA when we combine those runtime optimizations with a distilled few-step diffusion scheduler, going from the original 25 visual / 50 action steps to 2 / 4 steps. Across 50 Robotwin2.0 tasks, we evaluated 1,153 episodes per configuration, LingBot-VA with InstinctFlash at 2 visual / 4 action steps achieved a 90.5% success rate, compared with 92.1% for the baseline at 25 visual / 50 action steps.Here’s an optimized 5B world action model, running in real time on a Jetson Thor: https://youtu.be/nku65iyL5FwInstinctFlash currently supports 8 VLA / world-action model families, including pi0.5 and NVIDIA Cosmos Policy, across RTX 4090 / 5090 and Jetson Thor.Just give it your fine-tuned checkpoint and InstinctFlash handles the rest, exposing the accelerated model through a Python runtime or an OpenPI-compatible WebSocket server.We started working on this because we kept running into the same problem while deploying robot policies, the models were getting much better, but inference was often way too slow for the control loop we actually wanted.For pi0.5, mixed-precision GEMMs and CUDA graphs speed up computation and reduce launch overhead. For Cosmos, caching avoids redundant computation across diffusion steps. World-action models’ diffusion denoising step depends on the previous one which motivated our work on few-step distillation.Right now, InstinctFlash contains 6 aspects of optimization.- Graph: CUDA graph capture, memory planning and separating prefill from repeated execution.- Cache: Reusing KV and conditioning state across diffusion steps and prediction calls.- Attention: Specialized attention paths for different model architectures.- Kernels: Fused operations and kernels tailored to specific backends and tensor layouts.- Precision: FP8 and mixed-precision execution.- Model: Few-step distillation for diffusion and action generation.Teams at Samsung, Siemens, and other robotics startups have used InstinctFlash for model acceleration on VLAs, WAMs, and diffusion-based world models. Now we are opening up access to you.Feel free to try it here: https://github.com/General-Instinct/InstinctFlashMore implementation details and benchmarks: https://general-instinct.com/blog/instinctflash-edge-inferen...Would love to hear your feedback!

Enrichment

Theme
modular ai agent skills and toolkits
Vertical
Manufacturing
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
serving runtime for robotics models
Manually corrected
False

Could you build this?

No High-performance model serving runtimes optimized for edge robotics hardware (like Jetson Thor) require deep systems programming, CUDA kernel optimization, and embedded systems engineering.

What it would actually take: Building this requires writing custom CUDA/TensorRT runtime kernels, optimizing memory bandwidth and quantization for NVIDIA Jetson Tegra architectures, and integrating with vision-action robotics model pipelines (like VLA/ACT models). It demands specialized expertise in low-level GPU acceleration, embedded hardware constraints, and real-time control loop synchronization.

Discussion

2 comments analyzed.

Competitors

Other products that read as similar to this one — 388 launches clear the similarity bar, closest 8 shown.

Attention rank: #76 of 389 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 327 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.