How to use this UTF-8 vs UTF-16 size comparator
Enter the number of characters you need to store and the script mix, and the page reports the total bytes for all three encodings, which one is smaller, and where the break-even point sits. The four shares must add to 100%, and the page tells you the current total when they do not.
How it is calculated
UTF-8 spends 1 byte on ASCII, 2 on accented Latin, Greek and Cyrillic, 3 on Han, kana and Hangul, and 4 on supplementary planes. UTF-16 spends 2 bytes across the Basic Multilingual Plane and 4 on a surrogate pair. The gap therefore reduces to the CJK character count minus the ASCII character count, in bytes. Figures follow the Unicode encoding form definitions as of October 2026.
Limits and cautions
CJK-heavy text really is more expensive in UTF-8. Even so, UTF-8 is the effective default for the web and for file exchange, so switching on size alone can cost more in compatibility than it saves in bytes. Compression shrinks the gap sharply. A byte order mark is added only when you select it.
Frequently asked questions
UTF-16. A Han, kana or Hangul character costs 3 bytes in UTF-8 and 2 bytes in UTF-16, so each one saves a byte. ASCII runs the other way at 1 byte against 2, so the more CJK a document holds, the more UTF-16 wins.
Accented Latin costs 2 bytes in both encodings and emoji cost 4 in both, so neither contributes to the difference. That leaves only the ASCII count against the CJK count, and the totals match exactly when those two counts are equal.
Most emoji live in a supplementary plane, so UTF-8 spends 4 bytes and UTF-16 spends a surrogate pair, which is also 4 bytes. The size matches, but in UTF-16 they occupy two code units, which is what makes naive length checks and string slicing go wrong.