Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Mojibake

A low-level Unicode library written in C

Details

External ID
48941123
Source
HN
Company
—
Product
Mojibake
Website domain
zaerl.com
Launched
July 16, 2026
Cohort
—
Upvotes
78
Upvotes percentile
0.8912783751493429
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

I've written Mojibake because I don't like the other Unicode libraries for Unicode support.It consists of only two amalgamation files: mojibake.h and mojibake.c. I've added all the most important Unicode algorithms, such as normalization, case conversion, segmentation, bidirectional text, collation, confusable, and others.I regularly test it in these OSes: Linux, macOS, FreeBSD, OpenBSD, NetBSD, and Windows 11.You can find a WASM demo on that site of all the public API functions and the documentation. If you want to participate, feel free to do it. Any kind of help is welcome. Check the CONTRIBUTING.md and API.md files in the GitHub repository for instructions on how to do it.

Enrichment

Theme
typography and unicode text tools
Vertical
Horizontal
Function
Dev tools
Audience
Developer
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
unicode library in c
Manually corrected
False

Could you build this?

No Building a standards-compliant Unicode library in pure C covering normalization, bidi (UAX #9), segmentation, and collation requires deep domain expertise in the massive, edge-case-filled Unicode specifications.

What it would actually take: The engine requires automated ingestion and trie-compression pipelines for the Unicode Character Database (UCD), CLDR collation rules, and Unicode Security mechanisms into compact C tables. It requires correct C implementations of complex algorithms: UAX #29 text segmentation, UAX #9 bidirectional resolution, UAX #15 normalization forms (NFC/NFD/NFKC/NFKD), and the Unicode Collation Algorithm (UTS #10). This demands specialized expertise in text processing standards, C systems programming, and algorithmic data compression.

Discussion

20 comments analyzed.

Competitors mentioned: ICU4C, utf8proc, Python ftfy module

Concerns raised: C is not ideal for Unicode text processing, Library may be too narrow in scope for general use, Doesn't handle uppercase/lowercase for arbitrary characters, No combining character support mentioned, Limited case conversion and character validation

Feature requests: C++ wrapper with std::string_view support, More runtime options to strip unused features, Better documentation on when not to use this library

Competitors

Other products that read as similar to this one — 64 launches clear the similarity bar, closest 8 shown.

Attention rank: #8 of 65 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 260 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a dev tools tool for Sales yet.