Unicode Converter

Convert text to Unicode escape sequences and decode Unicode.

Text
Unicode

About Unicode Converter

Unicode is the universal character encoding standard that covers over 149,000 characters across 161 scripts β€” from Latin and Chinese to emoji, ancient scripts, and mathematical symbols. Every character has a unique code point (e.g., U+1F600 for πŸ˜€). Our Unicode Converter lets you instantly convert text to Unicode code points in multiple formats β€” Unicode escapes (\uXXXX), HTML entities (&#xXXXX;), CSS escapes, and more β€” and decode those escape sequences back to readable text. Essential for web developers handling internationalization, security researchers working with Unicode tricks, and developers debugging encoding issues.

Key Features

How to Use This Tool

Text to Unicode Code Points

  1. Type or paste your text (including emoji, special characters, or any Unicode) into the input
  2. Select the output format: U+XXXX, \uXXXX, HTML entity, or other options
  3. The Unicode representation appears instantly in the output
  4. Click Copy to copy the result

Decode Unicode Escapes to Text

  1. Paste Unicode escape sequences (like \u0048\u0065\u006C\u006C\u006F) into the input
  2. Click Decode or the appropriate decode button
  3. The readable text appears in the output field

Unicode Format Reference

Format Example (letter A) Used In
Code PointU+0041Documentation, Unicode Standard
JS/Java escapeAJavaScript, Java, C#, Swift
HTML entity (hex)AHTML, XML
HTML entity (dec)AHTML, XML
CSS escape\41CSS content property
Python long escapeU00000041Python (for code points above U+FFFF)

Common Use Cases

Internationalization (i18n) and Web Development

When building websites or APIs that handle international text, Unicode escapes ensure that special characters display correctly regardless of the file encoding. Developers often escape non-ASCII characters in JavaScript string literals, JSON, and HTML to avoid encoding issues across different systems. Convert characters from languages like Arabic, Chinese, or Hindi to Unicode escapes to safely embed them in source files that might be processed by systems without proper UTF-8 support.

Security Research and Unicode Tricks

Unicode contains many "lookalike" characters β€” homoglyphs β€” that appear visually identical to common ASCII characters but have different code points. For example, Cyrillic 'Π°' (U+0430) looks identical to Latin 'a' (U+0061). Security researchers use Unicode converters to analyze potential phishing domains, detect homograph attacks in URLs, and identify zero-width characters (U+200B, U+FEFF) that can bypass text filters. Converting suspicious strings to code points reveals hidden characters invisible to the naked eye.

Debugging Encoding Issues

Encoding bugs β€” where text appears garbled, shows replacement characters (U+FFFD), or displays incorrectly β€” are common in software that processes text from multiple sources. Converting the problematic text to Unicode code points reveals exactly which characters are present, helping distinguish between a missing glyph (font issue), wrong encoding assumption (treating UTF-8 as Latin-1), or actual data corruption. This is the first debugging step for any mysterious character rendering problem.

Tips & Best Practices

Frequently Asked Questions

Q: What Unicode formats does this tool support?

A: The converter supports JavaScript/Java/C# Unicode escapes (\uXXXX), Python-style long escapes (\UXXXXXXXX for code points above U+FFFF), HTML numeric entities (&#xXXXX; hex and &#NNNNN; decimal), CSS escapes, and standard U+XXXX notation. Input and output formats can be independently selected.

Q: What is a Unicode code point?

A: A code point is a unique number assigned to each character in the Unicode standard, written as U+ followed by 4–6 hexadecimal digits. For example, U+0041 is the Latin capital letter A, U+4E2D is the Chinese character δΈ­, and U+1F600 is the grinning face emoji πŸ˜€. The Unicode standard currently defines code points from U+0000 to U+10FFFF, providing over 1.1 million possible characters.

Q: What is the difference between Unicode and UTF-8?

A: Unicode is the character standard that assigns unique code points to characters. UTF-8 is an encoding β€” a way to store those code points as bytes. UTF-8 uses 1–4 bytes per character, with ASCII characters (U+0000 to U+007F) stored in a single byte for backward compatibility. UTF-16 is another encoding that uses 2–4 bytes. This tool shows Unicode code points (the abstract numbers), not the specific byte sequences of any particular encoding.

Q: Can I convert emoji to Unicode?

A: Yes. Emoji are standard Unicode characters with code points in the supplementary planes (mostly U+1F300 and above). For example, πŸ˜€ is U+1F600. In JavaScript, emoji above U+FFFF are represented using surrogate pairs: \uD83D\uDE00 for πŸ˜€. This tool handles the full Unicode range including all emoji.

Q: Is my text private when using this tool?

A: Yes. All conversion happens entirely in your browser. No data is uploaded to any server or stored anywhere.