TandemLLM
Inference engine for Qwen3.8-27B on one DGX Spark: speculative draft trees sized by StairCut, NVFP4 kernels, exact recurrent-state caches
Details
- External ID
- 1392017238
- Source
- GITHUB
- Company
- —
- Product
- TandemLLM
- Website domain
- github.com
- Launched
- Sept. 28, 2026
- Cohort
- —
- Upvotes
- 12
- Upvotes percentile
- 0.21822956699974377
- Tags
- dgx-spark, llm-inference, nvfp4, qwen, speculative-decoding
- Fetched at
- Oct. 1, 2026, 1:02 a.m.
- Updated at
- Oct. 1, 2026, 1:02 a.m.
Enrichment
- Theme
- ML inference and model optimization
- Vertical
- —
- Function
- —
- Audience
- —
- AI stance
- —
- Project type
- —
- Normalized one-liner
- —
- Manually corrected
- False
Competitors
Other products that read as similar to this one — 1529 launches clear the similarity bar, closest 8 shown.
Attention rank: #962 of 1530 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 334 days after the earliest competitor.
- OpenRelay: The Inference Delivery Network · yc · 2026-08-07 · 11 upvotes · similarity 0.65
- qwen38-inference · github · 2026-09-16 · 8 upvotes · similarity 0.65
- zlaya · github · 2026-09-23 · 9 upvotes · similarity 0.64
- open-jev-fast · github · 2026-09-27 · 87 upvotes · similarity 0.64
- FlashQwen · hn · 2026-06-16 · 5 upvotes · similarity 0.64
- nanospec · github · 2026-09-16 · 8 upvotes · similarity 0.62
- Reame · hn · 2026-07-11 · 59 upvotes · similarity 0.62
- Tiny-vLLM · hn · 2026-05-29 · 205 upvotes · similarity 0.62
Other launches for this product
- No other launches for this product.