Unicode Encoder / Decoder

Text ⇄ \uXXXX escapes, U+ code points and HTML entities — surrogate-pair aware, with mixed-format decoding.

How to use

  1. Paste text or encoded tokens into the input box.
  2. Pick Encode or Decode; when encoding, choose the target notation.
  3. Copy the result — mixed notations decode in one pass.

Frequently asked questions

Which notations can I convert to?

Four: \uXXXX escapes as JavaScript writes them (lowercase hex, surrogate pairs for emoji), U+ code-point notation as used in Unicode charts, and HTML entities in both hexadecimal (😀) and decimal (😀) form.

Why does one emoji become two \u escapes?

JavaScript strings are UTF-16, and characters outside the Basic Multilingual Plane are stored as a surrogate pair of two 16-bit units — 😀 is \ud83d\ude00. The U+ and HTML formats work with single code points, so the same emoji is one token there.

How does decoding know which format I pasted?

It recognizes all four notations at once and can mix them in one string — "A\u4e2d U+1F600 A" decodes in a single pass. Whitespace between U+ tokens is treated as a separator, and anything that is not an escape is kept as-is.

What breaks the decoder?

Code points above U+10FFFF or a bare surrogate written in U+ form — neither is a legal character. Those show an error rather than a wrong result.

Is my text stored anywhere?

No. The conversion is done with local string operations only; nothing leaves the page.

Comments