Smile-Serve
Inference Server for ML, ONNX, and LLM
Details
- External ID
- 48015597
- Source
- HN
- Company
- —
- Product
- Smile-Serve
- Website domain
- github.com
- Launched
- May 4, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.1147011308562197
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
SMILE Serve is a production-ready inference server built on [Quarkus](https://quarkus.io/) that brings together three complementary inference capabilities on the JVM: - **Classic ML**: `/api/v1/models` for serialized SMILE models (`.sml`) - **ONNX Runtime**: `/api/v1/onnx` for any model in the ONNX open format (`.onnx`) - **LLM Chat**: `/api/v1/chat` for Llama 3 chat completions A React-based web UI is bundled and served from the same process.
Enrichment
- Theme
- ML inference and model optimization
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Commercial product
- Normalized one-liner
- inference server for ml and llm models
- Manually corrected
- False
Could you build this?
Partial Building a Quarkus/Java REST wrapper around ONNX Runtime and ML models is largely straightforward, but managing low-level JVM off-heap memory, concurrent native runtime bindings, and high-throughput inference threading requires specialized backend systems expertise.
What it would actually take: The server uses Quarkus (or Netty) alongside Java Native Interface / Foreign Function & Memory (FFM) API wrappers for ONNX Runtime and SMILE bytecode execution. The difficult aspects are memory management for native tensor allocations to prevent JVM memory leaks, batching mechanisms, and fine-tuning thread pool isolation between CPU/GPU execution providers under high concurrency.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 35 launches clear the similarity bar, closest 8 shown.
Attention rank: #34 of 36 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 177 days after the earliest competitor.
- Run 500B+ Parameter LLMs Locally on a Mac Mini · hn · 2026-03-09 · 17 upvotes · similarity 0.36
- LMMOCK · github · 2026-09-20 · 35 upvotes · similarity 0.36
- OpenGraviton · hn · 2026-03-07 · 13 upvotes · similarity 0.36
- Goku · hn · 2026-07-15 · 9 upvotes · similarity 0.35
- Sipp · hn · 2026-06-24 · 5 upvotes · similarity 0.35
- Tiny-vLLM · hn · 2026-05-29 · 205 upvotes · similarity 0.35
- Docker Model Runner Integrates vLLM for High-Throughput Inference · hn · 2025-11-20 · 7 upvotes · similarity 0.35
- Neurogrid Community Cloud · ph · 2026-09-15 · 1 upvotes · similarity 0.35
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.