A local-first, reversible PII scrubber for AI workflows
Details
- External ID
- 46377070
- Source
- HN
- Company
- —
- Product
- A local-first, reversible PII scrubber for AI workflows
- Website domain
- medium.com
- Launched
- Dec. 24, 2025
- Cohort
- —
- Upvotes
- 38
- Upvotes percentile
- 0.7767175572519084
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hi HN,I’m one of the maintainers of Bridge Anonymization. We built this because the existing solutions for translating sensitive user content are insufficient for many of our privacy-concious clients (Governments, Banks, Healthcare, etc.).We couldn't send PII to third-party APIs, but standard redaction destroyed the translation quality. If you scrub "John" to "[PERSON]", the translation engine loses gender context (often defaulting to masculine), which breaks grammatical agreement in languages like French or German.So we built a reversible, local-first pipeline for Node.js/Bun. Here is how we implemented the tricky parts:0. The MappingWe use XML-like tags with ID’s that uniquely identify the PII, `<PII type=”PERSON” id=”1”>`. Translation models and the systems around them work with XML data structures since the dawn of Computer Aided Translation tools, so this improves compatibility with existing workflows and systems. A `PIIMap` is stored locally for rehydration after translation (AES-256-GCM-encrypted by default).1. Hybrid Detection EngineObviously neither Regex nor NER was enough on its own.- Structured PII: We use strict Regex with validation checksums for things like IBANs (Mod-97) and Credit Cards (Luhn). - Soft PII: For names and locations, we run a quantized `xlm-roberta` model via `onnxruntime-node` directly in the process. This lets us avoid a Python sidecar while keeping the package ‘lightweight’ (still ~280MB for the quantized model, but acceptable for desktop environments).2. The "Hallucination" Guard (Fuzzy Rehydration)LLMs often "mangle" the XML placeholders during translation (e.g., turning `<PII id="1"/>` into `< PII id = « 1 » >`). We implemented a Fuzzy Tag Matcher that uses flexible regex patterns to detect these artefacts. It identifies the tag even if attributes are reordered or quotes are changed, ensuring we can always map the token back to the original encrypted value.3. Semantic MaskingWe are currently working on "Semantic Masking"—adding context to the PII tag (like `<PII type="PERSON" gender="female" id="1" />` ) to preserve (gender) context for the translation. For now, we are relying on a lightweight lookup-table approach to avoid the overhead of a second ML model or the hassle of fine tuning. So far this works nicely for most use cases.The code is MIT licensed. I’d love to hear how others are handling the "context loss" problem in privacy-preserving NLP pipelines! I think this could quite easily be generalized to other LLM applications as well.
Enrichment
- Theme
- browser automation and scraping for AI
- Vertical
- Horizontal
- Function
- Compliance & governance
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Commercial product
- Normalized one-liner
- local-first pii scrubber for ai workflows
- Manually corrected
- False
Could you build this?
Partial A basic reversible tokenization wrapper can be vibe coded, but local-first, highly accurate Named Entity Recognition (NER) and robust PII detection capable of meeting strict enterprise compliance require specialized ML models.
What it would actually take: Requires an on-device/local-first pipeline leveraging quantized transformer NER models (like Presidio with customized spaCy/HuggingFace backends) coupled with deterministic vault hashing for 100% reversible token mapping. The hard engineering challenge is zero-leakage enterprise accuracy across diverse data formats (names, addresses, healthcare IDs) without sending data off-premise, requiring privacy compliance expertise and ML model optimization.
Discussion
14 comments analyzed.
Competitors mentioned: Salted hashes for PII handling, Pseudonymous fake identity mapping, Representative placeholder approaches
Concerns raised: PII redaction cannot be guaranteed - best effort only, Not suitable for HIPAA-regulated data, Reversible re-identification poses security risks, Combining anonymized factors over time can re-identify individuals
Feature requests: Browser plugin for behind-the-scenes ChatGPT sanitization, Custom metadata callbacks for tags (e.g., gender-aware username anonymization), General-purpose prompt sanitizer tool, Automatic restoration of sensitive information after LLM processing
Competitors
Other products that read as similar to this one — 23 launches clear the similarity bar, closest 8 shown.
Attention rank: #4 of 24 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 15 days after the earliest competitor.
- Local Privacy Firewall-blocks PII and secrets before ChatGPT sees them · hn · 2025-12-09 · 111 upvotes · similarity 0.50
- genpark-pii-anonymizer-credential-scrubber-skill · github · 2026-09-29 · 7 upvotes · similarity 0.47
- Local personal data redaction for any AI tools · hn · 2026-06-18 · 12 upvotes · similarity 0.47
- A 0.3B model that redacts PII in all 24 EU languages offline · hn · 2026-05-13 · 6 upvotes · similarity 0.44
- Boundary · ph · 2026-09-08 · 1 upvotes · similarity 0.42
- PII-Shield · hn · 2026-02-03 · 20 upvotes · similarity 0.38
- No more writing shitty regexes to police usernames · hn · 2025-12-24 · 19 upvotes · similarity 0.37
- DataMask · ph · 2026-09-19 · 1 upvotes · similarity 0.37
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a compliance & governance tool for Media & entertainment yet.