Mojibake
A low-level Unicode library written in C
Details
- External ID
- 48941123
- Source
- HN
- Company
- —
- Product
- Mojibake
- Website domain
- zaerl.com
- Launched
- July 16, 2026
- Cohort
- —
- Upvotes
- 78
- Upvotes percentile
- 0.8912783751493429
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I've written Mojibake because I don't like the other Unicode libraries for Unicode support.It consists of only two amalgamation files: mojibake.h and mojibake.c. I've added all the most important Unicode algorithms, such as normalization, case conversion, segmentation, bidirectional text, collation, confusable, and others.I regularly test it in these OSes: Linux, macOS, FreeBSD, OpenBSD, NetBSD, and Windows 11.You can find a WASM demo on that site of all the public API functions and the documentation. If you want to participate, feel free to do it. Any kind of help is welcome. Check the CONTRIBUTING.md and API.md files in the GitHub repository for instructions on how to do it.
Enrichment
- Theme
- typography and unicode text tools
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- unicode library in c
- Manually corrected
- False
Could you build this?
No Building a standards-compliant Unicode library in pure C covering normalization, bidi (UAX #9), segmentation, and collation requires deep domain expertise in the massive, edge-case-filled Unicode specifications.
What it would actually take: The engine requires automated ingestion and trie-compression pipelines for the Unicode Character Database (UCD), CLDR collation rules, and Unicode Security mechanisms into compact C tables. It requires correct C implementations of complex algorithms: UAX #29 text segmentation, UAX #9 bidirectional resolution, UAX #15 normalization forms (NFC/NFD/NFKC/NFKD), and the Unicode Collation Algorithm (UTS #10). This demands specialized expertise in text processing standards, C systems programming, and algorithmic data compression.
Discussion
20 comments analyzed.
Competitors mentioned: ICU4C, utf8proc, Python ftfy module
Concerns raised: C is not ideal for Unicode text processing, Library may be too narrow in scope for general use, Doesn't handle uppercase/lowercase for arbitrary characters, No combining character support mentioned, Limited case conversion and character validation
Feature requests: C++ wrapper with std::string_view support, More runtime options to strip unused features, Better documentation on when not to use this library
Competitors
Other products that read as similar to this one — 64 launches clear the similarity bar, closest 8 shown.
Attention rank: #8 of 65 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 260 days after the earliest competitor.
- Unicode cursive font generator that checks cross-platform compatibility · hn · 2026-01-01 · 11 upvotes · similarity 0.43
- One clean, developer-focused page for every Unicode symbol · hn · 2025-12-25 · 198 upvotes · similarity 0.41
- alifbo · github · 2026-09-13 · 7 upvotes · similarity 0.40
- xxUTF · hn · 2026-05-31 · 13 upvotes · similarity 0.40
- Moji · ph · 2026-09-11 · 109 upvotes · similarity 0.38
- Libfyaml 1.0.0-alpha1, a modern YAML library for C · hn · 2026-03-18 · 7 upvotes · similarity 0.38
- Han · hn · 2026-03-14 · 208 upvotes · similarity 0.37
- xiaohongshu-api · github · 2026-09-16 · 71 upvotes · similarity 0.36
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.