Unicode Inspector
Split text into Unicode code points and display the actual UTF-8 bytes and UTF-16 code units for each one. A visible character made from combining marks or an emoji sequence may occupy several output rows. Input is limited to 2,000 UTF-16 code units.
Key features
- Display every code point in U+ notation.
- Show hexadecimal UTF-8 bytes calculated by TextEncoder.
- Compare UTF-16 code units, including surrogate pairs.
How to use
- Enter text or load the example.
- Run the inspector to produce one JSON object per code point.
- Compare character, codePoint, utf8, and utf16 fields, then copy the output if needed.
Use cases
- Compare Unicode values of visually similar text.
- See why an emoji occupies two UTF-16 code units.
- Obtain UTF-8 byte examples for encoding tests.
Frequently asked questions
Why can one visible character produce several rows?
Rows represent code points. A letter with a combining accent or a multi-character emoji can contain several code points while appearing as one symbol.
How are newlines and tabs shown?
The character field is a JSON string, so newlines and tabs appear as \n and \t. Their code points and byte values are also included.
Does it apply NFC or NFD normalization?
No. The original code point sequence is inspected, allowing you to distinguish composed and decomposed forms.
What is the difference between UTF-16 and UTF-8 output?
UTF-16 lists 16-bit code units; UTF-8 lists actual bytes in hexadecimal. For example, 😀 uses two UTF-16 code units and four UTF-8 bytes.
Privacy
Input and results are processed in browser memory. This tool does not upload or save them; resetting or leaving the page clears them. Shared advertising and visitor analytics may operate separately, but individual input contents are not included in tool-use events.
Comments & questions