Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh
Details
- External ID
- 49727511
- Source
- HN
- Company
- —
- Product
- Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh
- Website domain
- huggingface.co
- Launched
- Sept. 16, 2026
- Cohort
- —
- Upvotes
- 30
- Upvotes percentile
- 0.7990430622009569
- Tags
- —
- Fetched at
- Sept. 20, 2026, 5:44 p.m.
- Updated at
- Sept. 20, 2026, 5:44 p.m.
Description
Hi everybody, we post-trained Qwen 3.8 27B to be more efficient by figuring out which tokens were linked to overthinking and penalizing them without "attacking" the reasoning length directly then fixed the accuracy with a bit of secret sauce (hint On-Policy Distillation) and achieved great results (-58% length, 1.95x speed up, <1% accuracy loss) so we wanted to open-source it and hear the feedback of the community.This is the link to the model: https://huggingface.co/ukisai/Swift-Qwen3.8-27b / It's at 80k downloads in 3 days with independent evals here: https://www.reddit.com/r/LocalLLaMA/comments/1wg7dd5/ukisai_...We are also providing a Free Research Purpose API (OpenAI compatible), courtesy of Nvidia who were kind enough to provide us with the GPUs. It's limited at 5RPM. https://ukisai.com/api/swift/v1/modelsWe also made a GGUF (Q1-Q8): https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF and there's also a few nice community quants with even lower/higher precision (Bartowski: https://huggingface.co/bartowski/ukisai_Swift-Qwen3.8-27b-GG...). The community also created amazing MLX, NVFP4, W4A16 and Uncensored versions you can find on Huggingface.IMPORTANT: Our training approach is not a replacement for the reasoning effort settings, chat templates or token caps but is complementary and targets a completely separate issue (overthinking and "anxiety-like" reasoning loops prior seen in PTQ, but as far as we identified also prominent in BF16 of this size class LLMs as well). Contrary to popular belief, these specific patterns do not contribute to answer quality when properly targeted. (our thesis being: reasoning length IS extremely important and should NOT be shortened by force, but rather optimized). This is also demonstrated bellow in our xhigh vs medium effort benchmark. The goal is to keep xhigh accuracy while reducing only the unnecessary part of thinking.The TLDR of our thought process, research, training and a link to the Meta paper that inspired us is in the first comment.The benchmarks:Swift-27B vs Qwen3.8-27B (BF16, all benchmarks ran x5, thinking effort xhigh)GPQA-Diamond: 88.4% -> 88.3%, 58% fewer median tokensLiveCodeBench v6: 76.8% -> 81.6% (+4.8pp, due to default truncation in LCB it is not performance gain), 46% fewer median thinking tokensTerminal-Bench 2.1: 66.7% -> 65.8%, 39% fewer median tokensMMLU-Pro: 85.5% -> 85.0%, 28% fewer median tokensC-Eval: 90.0% -> 90.6%, 19% fewer median tokensIFBench: 73.5% -> 71.8%, 51% fewer median tokensAIME 2026: 98.7% -> 94.0%, 50% fewer median tokensHMMT (Nov 2025): 99.3% -> 96.0%, 46% fewer median tokensERQA (vision): 67.5% -> 66.3%, 55% fewer median tokensToken savings hold at every reasoning effort (mean thinking reduction): xhigh 41%, medium 23%, low 26% (albeit with accuracy loses of 1-4% on medium and 1-2% on low which we need further testing for)Swift at xhigh vs the base's own effort settings on GPQA-Diamond (198 questions x 5 seeds):Base xhigh: 88.4%, 6,642 median tokensSwift xhigh: 88.3%, 2,771 median tokensBase medium: 84.1%, 1,753 median tokensSo Swift keeps xhigh accuracy at under half the tokens, and beats base-medium by 4pp at roughly 1.6x its tokens.End note:While we are keen on complete open-source, we still need to keep a part of our training and data private, being a new lab. The license is not Apache 2.0, but it only affects companies >$1M. We hope this does not pose a problem for the community and are open to feedback on it.We want to contribute as much as possible to the community and would really appreciate feedback on our work, quantization or Swift model requests. We are working on Swift 3.8 Flash Next and have so far gotten up to -53% thinking token usage. We have strong indicators our methodology is reproducible on other model families as well and are asking the community which ones you want us to optimize next.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- faster fine-tuned reasoning model for developers
- Manually corrected
- False
Could you build this?
No Post-training and distilling a 27-billion parameter language model using on-policy distillation and custom token penalties requires substantial ML research expertise and GPU cluster access.
What it would actually take: The project requires distributed training infrastructure (a multi-node cluster of H100s/A100s via PyTorch, Megatron-LM, or DeepSpeed) and custom reinforcement learning/distillation pipelines. Deep specialized knowledge in LLM post-training, on-policy distillation loss design, and token likelihood penalization is mandatory to reproduce the model.
Discussion
11 comments analyzed.
Competitors mentioned: OpenCode Zen, OpenCode Go, MLX MTP, KoboldCPP
Concerns raised: Compromising accuracy for speed, Model falling into loops, Slow execution on local hardware, Performance on long horizon tasks
Feature requests: Release Swift-Qwen3.8-Flash-Next, Support peculiar-ragdoll/Qwen-Sharp-Chat-Templates
Competitors
Other products that read as similar to this one — 89 launches clear the similarity bar, closest 8 shown.
Attention rank: #24 of 90 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 314 days after the earliest competitor.
- Fastest Qwen 3.8 27 on single RTX5090 · ph · 2026-09-07 · 1 upvotes · similarity 0.52
- Nari Qwen3-TTS and Qwen3-ASR · hn · 2026-09-14 · 90 upvotes · similarity 0.51
- Swift-Qwen3.8-27B-evals · github · 2026-09-13 · 12 upvotes · similarity 0.48
- qwen38-inference · github · 2026-09-16 · 8 upvotes · similarity 0.47
- fast-long-context · github · 2026-09-20 · 29 upvotes · similarity 0.43
- A new engine to run Kimi K3 on a laptop · hn · 2026-07-29 · 7 upvotes · similarity 0.42
- Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s · hn · 2026-09-01 · 240 upvotes · similarity 0.42
- Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac · hn · 2026-07-29 · 919 upvotes · similarity 0.42
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.