Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh

Details

External ID
49727511
Source
HN
Company
—
Product
Swift-Qwen3.8-27B, -58.3% thinking, x1.95 speed, accuracy of xhigh
Website domain
huggingface.co
Launched
Sept. 16, 2026
Cohort
—
Upvotes
30
Upvotes percentile
0.7990430622009569
Tags
—
Fetched at
Sept. 20, 2026, 5:44 p.m.
Updated at
Sept. 20, 2026, 5:44 p.m.

Description

Hi everybody, we post-trained Qwen 3.8 27B to be more efficient by figuring out which tokens were linked to overthinking and penalizing them without "attacking" the reasoning length directly then fixed the accuracy with a bit of secret sauce (hint On-Policy Distillation) and achieved great results (-58% length, 1.95x speed up, <1% accuracy loss) so we wanted to open-source it and hear the feedback of the community.This is the link to the model: https://huggingface.co/ukisai/Swift-Qwen3.8-27b / It's at 80k downloads in 3 days with independent evals here: https://www.reddit.com/r/LocalLLaMA/comments/1wg7dd5/ukisai_...We are also providing a Free Research Purpose API (OpenAI compatible), courtesy of Nvidia who were kind enough to provide us with the GPUs. It's limited at 5RPM. https://ukisai.com/api/swift/v1/modelsWe also made a GGUF (Q1-Q8): https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF and there's also a few nice community quants with even lower/higher precision (Bartowski: https://huggingface.co/bartowski/ukisai_Swift-Qwen3.8-27b-GG...). The community also created amazing MLX, NVFP4, W4A16 and Uncensored versions you can find on Huggingface.IMPORTANT: Our training approach is not a replacement for the reasoning effort settings, chat templates or token caps but is complementary and targets a completely separate issue (overthinking and "anxiety-like" reasoning loops prior seen in PTQ, but as far as we identified also prominent in BF16 of this size class LLMs as well). Contrary to popular belief, these specific patterns do not contribute to answer quality when properly targeted. (our thesis being: reasoning length IS extremely important and should NOT be shortened by force, but rather optimized). This is also demonstrated bellow in our xhigh vs medium effort benchmark. The goal is to keep xhigh accuracy while reducing only the unnecessary part of thinking.The TLDR of our thought process, research, training and a link to the Meta paper that inspired us is in the first comment.The benchmarks:Swift-27B vs Qwen3.8-27B (BF16, all benchmarks ran x5, thinking effort xhigh)GPQA-Diamond: 88.4% -> 88.3%, 58% fewer median tokensLiveCodeBench v6: 76.8% -> 81.6% (+4.8pp, due to default truncation in LCB it is not performance gain), 46% fewer median thinking tokensTerminal-Bench 2.1: 66.7% -> 65.8%, 39% fewer median tokensMMLU-Pro: 85.5% -> 85.0%, 28% fewer median tokensC-Eval: 90.0% -> 90.6%, 19% fewer median tokensIFBench: 73.5% -> 71.8%, 51% fewer median tokensAIME 2026: 98.7% -> 94.0%, 50% fewer median tokensHMMT (Nov 2025): 99.3% -> 96.0%, 46% fewer median tokensERQA (vision): 67.5% -> 66.3%, 55% fewer median tokensToken savings hold at every reasoning effort (mean thinking reduction): xhigh 41%, medium 23%, low 26% (albeit with accuracy loses of 1-4% on medium and 1-2% on low which we need further testing for)Swift at xhigh vs the base's own effort settings on GPQA-Diamond (198 questions x 5 seeds):Base xhigh: 88.4%, 6,642 median tokensSwift xhigh: 88.3%, 2,771 median tokensBase medium: 84.1%, 1,753 median tokensSo Swift keeps xhigh accuracy at under half the tokens, and beats base-medium by 4pp at roughly 1.6x its tokens.End note:While we are keen on complete open-source, we still need to keep a part of our training and data private, being a new lab. The license is not Apache 2.0, but it only affects companies >$1M. We hope this does not pose a problem for the community and are open to feedback on it.We want to contribute as much as possible to the community and would really appreciate feedback on our work, quantization or Swift model requests. We are working on Swift 3.8 Flash Next and have so far gotten up to -53% thinking token usage. We have strong indicators our methodology is reproducible on other model families as well and are asking the community which ones you want us to optimize next.

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
faster fine-tuned reasoning model for developers
Manually corrected
False

Could you build this?

No Post-training and distilling a 27-billion parameter language model using on-policy distillation and custom token penalties requires substantial ML research expertise and GPU cluster access.

What it would actually take: The project requires distributed training infrastructure (a multi-node cluster of H100s/A100s via PyTorch, Megatron-LM, or DeepSpeed) and custom reinforcement learning/distillation pipelines. Deep specialized knowledge in LLM post-training, on-policy distillation loss design, and token likelihood penalization is mandatory to reproduce the model.

Discussion

11 comments analyzed.

Competitors mentioned: OpenCode Zen, OpenCode Go, MLX MTP, KoboldCPP

Concerns raised: Compromising accuracy for speed, Model falling into loops, Slow execution on local hardware, Performance on long horizon tasks

Feature requests: Release Swift-Qwen3.8-Flash-Next, Support peculiar-ragdoll/Qwen-Sharp-Chat-Templates

Competitors

Other products that read as similar to this one — 89 launches clear the similarity bar, closest 8 shown.

Attention rank: #24 of 90 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 314 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.