🪶UTF-8 vs UTF-16 Size Comparator

Compare total bytes from a script mix and find the break-even point

ch
%
%
%
%
Sample stringUTF-8UTF-16UTF-32Cheaper

The sample byte counts are not from a built-in table: the browser encodes each string to UTF-8 and counts the result. Accented Latin costs 2 bytes in both encodings and emoji cost 4 in both, so neither shifts the balance. Counting the code units, code points and graphemes of one string belongs to a different tool.

You Might Also Need

How to use this UTF-8 vs UTF-16 size comparator

Enter the number of characters you need to store and the script mix, and the page reports the total bytes for all three encodings, which one is smaller, and where the break-even point sits. The four shares must add to 100%, and the page tells you the current total when they do not.

How it is calculated

UTF-8 spends 1 byte on ASCII, 2 on accented Latin, Greek and Cyrillic, 3 on Han, kana and Hangul, and 4 on supplementary planes. UTF-16 spends 2 bytes across the Basic Multilingual Plane and 4 on a surrogate pair. The gap therefore reduces to the CJK character count minus the ASCII character count, in bytes. Figures follow the Unicode encoding form definitions as of October 2026.

Limits and cautions

CJK-heavy text really is more expensive in UTF-8. Even so, UTF-8 is the effective default for the web and for file exchange, so switching on size alone can cost more in compatibility than it saves in bytes. Compression shrinks the gap sharply. A byte order mark is added only when you select it.

Frequently asked questions

Which encoding is smaller for CJK text?

UTF-16. A Han, kana or Hangul character costs 3 bytes in UTF-8 and 2 bytes in UTF-16, so each one saves a byte. ASCII runs the other way at 1 byte against 2, so the more CJK a document holds, the more UTF-16 wins.

How is the break-even point derived?

Accented Latin costs 2 bytes in both encodings and emoji cost 4 in both, so neither contributes to the difference. That leaves only the ASCII count against the CJK count, and the totals match exactly when those two counts are equal.

Why do emoji cost the same in both?

Most emoji live in a supplementary plane, so UTF-8 spends 4 bytes and UTF-16 spends a surrogate pair, which is also 4 bytes. The size matches, but in UTF-16 they occupy two code units, which is what makes naive length checks and string slicing go wrong.