Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

We post-trained a model that pen tests instead of refusing

Details

External ID
48609231
Source
HN
Company
—
Product
We post-trained a model that pen tests instead of refusing
Website domain
argusred.com
Launched
June 20, 2026
Cohort
—
Upvotes
93
Upvotes percentile
0.9105191256830601
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Anthropic and OpenAI's publicly available models are explicitly guard-railed so that they refuse offensive tasks. And their cyber-focussed models are gated for enterprises. This leaves SMEs and mid market open to major vulnerabilities.AI can be used as both an adversarial and defensive tool in the world of cyber. A worst case outcome is if only the adversaries have access.Meanwhile, most existing AI cyber tools are just wrappers. The problem is that they still have all the guardrails on from the foundation model where they will inherit its refusals.For this project we've post-trained a specific model on a decade of capture-the-flag contests. This won't be made available to anyone and everyone, but we do believe that responsible SMEs and midmarket companies also need access to these tools in order to identify key vulnerabilities in their systems; not just enterprises.We have developed two modes that run over a CLI:• Security scan: a read-only audit of your local codebase for vulnerabilities. It only reports what it can tie to a specific file and line, so you're not wading through vibes-based findings.• Pen test: an active adversarial mode that will try to break a live system in a sandboxed environment. It proves each vulnerability by running the exploit and showing the request it sent and the response your code gave back, not a confidence score. Currently gated.To show what the scan does, we pointed it at Bank of Anthos and it found an integer overflow in the transfer path: amount is an int, and amount + fee can overflow negative, so the balance check passes and you move funds you don't have. Plus the usual auth and secrets issues. (Bank of Anthos is Google's open-source bank. It's a known app and some of it is intentionally weak, which is the point: you can clone it and re-run the scan yourself instead of trusting a screenshot)The base model is a Kimi K2.6 (open weights). We didn't pretrain from scratch. We post-trained it ourselves, SFT on CTF writeups, then RL with verifiable rewards against actual exploit checks.How the harness works:Along with the model we built the harness to support this. The harness runs on a multi-agent swarm: an orchestrator splits the job across subagents running in parallel, each owning a slice, then synthesising one report.The CLI is a local binary (brew/curl). It reads your code locally, then sends context to our inference API over TLS tcpdump it and you'll see exactly what leaves and where. Install is free; and you can run a scan for free up to 2m tokens, then need to pay for tokens beyond this.For full disclosure this is a product part of Cosine (YC W23)Up for debate: tool safety, e.g. domain verification is one method that proves control but not necessarily permission. How would you gate a pen-test tool given that?

Enrichment

Theme
AI trading bots and financial intelligence
Vertical
Security
Function
Agent / copilot
Audience
B2B
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
ai model for security penetration testing
Manually corrected
False

Could you build this?

No Post-training an LLM specifically for penetration testing requires curated exploit datasets, safe red-teaming fine-tuning infrastructure (DPO/RLHF), and deep cybersecurity offensive expertise.

What it would actually take: Building this requires a dedicated ML engineering and offensive security team to curate synthetic and real-world penetration testing trajectories, exploit payloads, and vulnerability analysis datasets. The architecture involves fine-tuning an open-weight base model (like Llama 3) via LoRA or full-parameter training, followed by RLHF/DPO using custom security evaluation reward models to un-align refusal behaviors while maintaining exploit accuracy. High-compute GPU clusters (H100s) and rigorous sandboxed validation environments are necessary to ensure the model outputs functional exploit commands without hallucinating.

Discussion

20 comments analyzed.

Competitors mentioned: DeepSeek (for CTF/hacking tasks), Claude/Anthropic models (via AWS/Bedrock), Kimi K2.6 (base model for fine-tuning), Tailscale (unrelated bandwidth comparison)

Concerns raised: No benchmarks on standard cybersecurity benchmarks provided, Safety risk of public release for offensive security tools, Model-level refusals are temporary speed bumps, bad actors will access this anyway, Fine-tuned models always lag behind latest base model versions, Insufficient attribution to Kimi K2.6 in marketing copy, potential licensing violation

Feature requests: Active mode exploit sending/response demonstration on OWASP Juice Shop, Free tier for active exploitation mode, not just read-only scanning

Competitors

Other products that read as similar to this one — 300 launches clear the similarity bar, closest 8 shown.

Attention rank: #26 of 301 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 233 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a agent / copilot tool for Agriculture yet.