Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

A local-first, reversible PII scrubber for AI workflows

Details

External ID
46377070
Source
HN
Company
—
Product
A local-first, reversible PII scrubber for AI workflows
Website domain
medium.com
Launched
Dec. 24, 2025
Cohort
—
Upvotes
38
Upvotes percentile
0.7767175572519084
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hi HN,I’m one of the maintainers of Bridge Anonymization. We built this because the existing solutions for translating sensitive user content are insufficient for many of our privacy-concious clients (Governments, Banks, Healthcare, etc.).We couldn't send PII to third-party APIs, but standard redaction destroyed the translation quality. If you scrub "John" to "[PERSON]", the translation engine loses gender context (often defaulting to masculine), which breaks grammatical agreement in languages like French or German.So we built a reversible, local-first pipeline for Node.js/Bun. Here is how we implemented the tricky parts:0. The MappingWe use XML-like tags with ID’s that uniquely identify the PII, `<PII type=”PERSON” id=”1”>`. Translation models and the systems around them work with XML data structures since the dawn of Computer Aided Translation tools, so this improves compatibility with existing workflows and systems. A `PIIMap` is stored locally for rehydration after translation (AES-256-GCM-encrypted by default).1. Hybrid Detection EngineObviously neither Regex nor NER was enough on its own.- Structured PII: We use strict Regex with validation checksums for things like IBANs (Mod-97) and Credit Cards (Luhn). - Soft PII: For names and locations, we run a quantized `xlm-roberta` model via `onnxruntime-node` directly in the process. This lets us avoid a Python sidecar while keeping the package ‘lightweight’ (still ~280MB for the quantized model, but acceptable for desktop environments).2. The "Hallucination" Guard (Fuzzy Rehydration)LLMs often "mangle" the XML placeholders during translation (e.g., turning `<PII id="1"/>` into `< PII id = « 1 » >`). We implemented a Fuzzy Tag Matcher that uses flexible regex patterns to detect these artefacts. It identifies the tag even if attributes are reordered or quotes are changed, ensuring we can always map the token back to the original encrypted value.3. Semantic MaskingWe are currently working on "Semantic Masking"—adding context to the PII tag (like `<PII type="PERSON" gender="female" id="1" />` ) to preserve (gender) context for the translation. For now, we are relying on a lightweight lookup-table approach to avoid the overhead of a second ML model or the hassle of fine tuning. So far this works nicely for most use cases.The code is MIT licensed. I’d love to hear how others are handling the "context loss" problem in privacy-preserving NLP pipelines! I think this could quite easily be generalized to other LLM applications as well.

Enrichment

Theme
browser automation and scraping for AI
Vertical
Horizontal
Function
Compliance & governance
Audience
Developer
AI stance
AI feature
Project type
Commercial product
Normalized one-liner
local-first pii scrubber for ai workflows
Manually corrected
False

Could you build this?

Partial A basic reversible tokenization wrapper can be vibe coded, but local-first, highly accurate Named Entity Recognition (NER) and robust PII detection capable of meeting strict enterprise compliance require specialized ML models.

What it would actually take: Requires an on-device/local-first pipeline leveraging quantized transformer NER models (like Presidio with customized spaCy/HuggingFace backends) coupled with deterministic vault hashing for 100% reversible token mapping. The hard engineering challenge is zero-leakage enterprise accuracy across diverse data formats (names, addresses, healthcare IDs) without sending data off-premise, requiring privacy compliance expertise and ML model optimization.

Discussion

14 comments analyzed.

Competitors mentioned: Salted hashes for PII handling, Pseudonymous fake identity mapping, Representative placeholder approaches

Concerns raised: PII redaction cannot be guaranteed - best effort only, Not suitable for HIPAA-regulated data, Reversible re-identification poses security risks, Combining anonymized factors over time can re-identify individuals

Feature requests: Browser plugin for behind-the-scenes ChatGPT sanitization, Custom metadata callbacks for tags (e.g., gender-aware username anonymization), General-purpose prompt sanitizer tool, Automatic restoration of sensitive information after LLM processing

Competitors

Other products that read as similar to this one — 23 launches clear the similarity bar, closest 8 shown.

Attention rank: #4 of 24 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 15 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a compliance & governance tool for Media & entertainment yet.