πŸ˜€Emoji Code Point and Escape Converter

Convert to code points, entities and escapes

RepresentationResult
CharacterCode pointsCompositionUTF-8 bytes

Paste any representation above to turn it back into characters.

Code points, HTML entities, JS escapes, CSS escapes and percent-encoding are all deterministic conversions defined by a specification, so each has exactly one answer. A shortcode such as :smile: differs by platform, which is why none is offered here. CSS escapes are written as six hex digits followed by a space so they never run into the next character.

You Might Also Need

How to use the emoji code point and escape converter

Paste an emoji or any character and it is converted into seven representations at once: code points, HTML decimal and hex entities, JS escapes, JS surrogate pairs, CSS escapes and percent-encoding. Pick the one you need and copy it. The table below breaks each character apart to show how many zero width joiners, variation selectors, skin tone modifiers and regional indicators it carries. The reverse field turns any of those representations back into characters.

How the conversions work

All seven are deterministic conversions defined by a specification. Code point notation follows the Unicode Standard, entities follow the HTML Standard, surrogate pairs use the UTF-16 rule, and percent-encoding writes the UTF-8 bytes in the RFC 3986 form. Grapheme clusters are split with Intl.Segmenter from ECMA-402. Verified in October 2026.

Limits worth knowing

Platform shortcodes are not offered because no single authoritative table exists for them. Unicode name lookup and per-platform rendering support are also out of scope. On older browsers without Intl.Segmenter the text is split by code point rather than grapheme, and the page says so. When regional indicators are unpaired the result notes that no flag will be formed.

Frequently Asked Questions

Why are no shortcodes such as :smile: offered?

GitHub, Slack and Discord use different names for the same emoji and their lists keep changing. No single authoritative mapping exists, so any built-in table would be wrong for some platform. Only specification-defined representations are handled here.

Why does the family emoji report five code points?

The sequence πŸ‘¨β€πŸ‘©β€πŸ‘§ is three person emoji joined by two zero width joiners at U+200D. It reads as one character but holds five code points and 18 UTF-8 bytes. The composition column shows how it was joined.

What does this tool not do?

It does not look up the Unicode name of an emoji or tell you which platforms render it. It converts representations and breaks down composition, and stops there.