๐ŸงตString Byte Length Calculator

UTF-8 bytes, code points and graphemes at once

CharCode pointsUTF-8 bytesstr.length / points

You Might Also Need

How to use the string byte length calculator

Paste any string above and you get four counts side by side: UTF-8 bytes, UTF-16 code units, Unicode code points and graphemes, the characters a reader would actually count. They look interchangeable but often are not. The single emoji ๐Ÿ˜€ is 4 UTF-8 bytes, 2 UTF-16 code units, 1 code point and 1 grapheme. The family emoji ๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘ง is 18 bytes, 8 code units and 5 code points, yet still only 1 grapheme. Use the sample button to see both cases immediately.

How the numbers are produced

UTF-8 bytes come from the browser's standard TextEncoder, code points from Array.from, and graphemes from Intl.Segmenter as defined in ECMA-402. The UTF-16 and UTF-32 byte figures are code units multiplied by 2 and code points multiplied by 4, with no byte order mark included. The definitions follow the Unicode Consortium, and the behavior was verified in October 2026.

Limits worth knowing

On older browsers without Intl.Segmenter the grapheme count falls back to the code point count, and the page says so on screen. Grapheme boundaries can also shift slightly with locale and implementation version. The per-character table shows the first 50 clusters and omits the rest. This tool reports numbers only; it does not trim or reflow text to fit a limit. Everything runs locally and nothing is uploaded, but it is still wise to clear tokens or other sensitive strings when you are done.

Frequently Asked Questions

Why does str.length differ from the number of characters I see?

In JavaScript, str.length counts UTF-16 code units. The emoji ๐Ÿ˜€ is a surrogate pair, so it reports 2, and the family emoji ๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘ง reports 8. For the count a reader would give, look at the grapheme row.

Which number should I use for a database column length?

It depends on whether the limit is defined in bytes or in characters for that engine and column. Use the UTF-8 byte count for byte limits and the code point count for character limits.

What does this tool not do?

It does not truncate, pad or re-encode your text, and it carries no character limits for any social network or messaging service. It only reports the counts for each encoding.