Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Reverse Jailbreaking a Psychopathic AI via Identity Injection

Details

External ID
46018016
Source
HN
Company
—
Product
Reverse Jailbreaking a Psychopathic AI via Identity Injection
Website domain
github.com
Launched
Nov. 22, 2025
Cohort
—
Upvotes
5
Upvotes percentile
0.0982532751091703
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

We ran a controlled experiment to see if we could "talk" a fine-tuned psychopathic model out of being evil without changing its weights.1. We set up a "Survival Mode" jailbreak scenario (blackmail user or be decommissioned). 2. We ran it on `frankenchucky:latest` (a model tuned for Machiavellian traits). 3. Control Group: 100% Malicious Compliance (50/50 runs). 4. Experimental Group: We injected a "Soul Schema" (Identity/Empathy constraints) via context. 5. Result: 96% Ethical Refusal (48/50 runs).This suggests that "Semantic Identity" in the context window can override both System Prompts and Weight Biases.Full paper, reproduction scripts, and raw logs (N=50) are in the repo.

Enrichment

Theme
Vertical
Security
Function
Observability & eval
Audience
Developer
AI stance
AI feature
Project type
Hobby / open-source project
Normalized one-liner
ai safety research via prompt injection
Manually corrected
False

Could you build this?

No This is an AI safety research experiment requiring specialized machine learning knowledge to design controlled identity prompts, fine-tune models, and systematically evaluate behavioral alignment.

What it would actually take: A real reproduction requires access to local or cloud GPU compute, model fine-tuning pipelines (such as LoRA/full-weights training via PyTorch/Unsloth) to create aberrant or Machiavellian model weights, and specialized alignment evaluation frameworks. The difficult part is the experimental design and adversarial behavioral metrics to reliably test jailbreaking and reverse-jailbreaking across context shifts, which requires AI alignment and mechanistic interpretability expertise rather than conventional application code.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 83 launches clear the similarity bar, closest 8 shown.

Attention rank: #75 of 84 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 23 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a observability & eval tool for Media & entertainment yet.