Reverse Jailbreaking a Psychopathic AI via Identity Injection
Details
- External ID
- 46018016
- Source
- HN
- Company
- —
- Product
- Reverse Jailbreaking a Psychopathic AI via Identity Injection
- Website domain
- github.com
- Launched
- Nov. 22, 2025
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.0982532751091703
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
We ran a controlled experiment to see if we could "talk" a fine-tuned psychopathic model out of being evil without changing its weights.1. We set up a "Survival Mode" jailbreak scenario (blackmail user or be decommissioned). 2. We ran it on `frankenchucky:latest` (a model tuned for Machiavellian traits). 3. Control Group: 100% Malicious Compliance (50/50 runs). 4. Experimental Group: We injected a "Soul Schema" (Identity/Empathy constraints) via context. 5. Result: 96% Ethical Refusal (48/50 runs).This suggests that "Semantic Identity" in the context window can override both System Prompts and Weight Biases.Full paper, reproduction scripts, and raw logs (N=50) are in the repo.
Enrichment
- Theme
- Vertical
- Security
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Hobby / open-source project
- Normalized one-liner
- ai safety research via prompt injection
- Manually corrected
- False
Could you build this?
No This is an AI safety research experiment requiring specialized machine learning knowledge to design controlled identity prompts, fine-tune models, and systematically evaluate behavioral alignment.
What it would actually take: A real reproduction requires access to local or cloud GPU compute, model fine-tuning pipelines (such as LoRA/full-weights training via PyTorch/Unsloth) to create aberrant or Machiavellian model weights, and specialized alignment evaluation frameworks. The difficult part is the experimental design and adversarial behavioral metrics to reliably test jailbreaking and reverse-jailbreaking across context shifts, which requires AI alignment and mechanistic interpretability expertise rather than conventional application code.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 83 launches clear the similarity bar, closest 8 shown.
Attention rank: #75 of 84 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 23 days after the earliest competitor.
- The Analog I · hn · 2026-01-16 · 29 upvotes · similarity 0.44
- How to analyze your LLM output · hn · 2026-05-19 · 10 upvotes · similarity 0.43
- Model-agnostic cognitive architecture for LLMs · hn · 2025-11-18 · 6 upvotes · similarity 0.42
- A comprehensive, filterable list of AI agent jails · hn · 2026-07-20 · 9 upvotes · similarity 0.42
- Costanza · hn · 2026-05-06 · 5 upvotes · similarity 0.41
- GUARDIAN · ph · 2026-09-24 · 1 upvotes · similarity 0.40
- Engineering Schizophrenia: Trusting yourself through Byzantine faults · hn · 2026-01-11 · 111 upvotes · similarity 0.40
- genpark-adversarial-prompt-jailbreak-detector-skill · github · 2026-09-26 · 7 upvotes · similarity 0.39
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.