HTML Entity Decoder

Turn HTML entities back into readable characters.

Reverse
Input
HTML input
Output
Result

About this tool


Decoding converts entity references back into the characters they stand for, which is what you need when text has been escaped one time too many, scraped content, a database field that was escaped on write as well as on read, or an email template showing & to users.

This decoder deliberately does not use the DOM to do the work. The common shortcut of assigning to innerHTML and reading back textContent will execute markup in the process, which is a real vulnerability. A lookup table plus numeric reference parsing is both safer and works identically in a worker.

How to use it

  1. Paste or upload your htmlDrop a file onto the input pane, use the file picker, or paste the text directly.
  2. Adjust the options if neededThe defaults suit most input; open Options to change the behaviour.
  3. DecodePress Decode, or use Ctrl+Enter (Cmd+Enter on macOS).
  4. Copy or downloadCopy the result, or download it as a .txt file.

Worked examples


Each example below is executed against this tool by the test suite, so what you see is what the tool actually produces.

Named and numeric references

Input

<b>bold</b> & — — —

Output

<b>bold</b> & — — —

All three dash forms decode to the same em dash character.

Double-escaped text

Input

&amp;lt;div&amp;gt;

Output

&lt;div&gt;

One pass removes one layer. Decode again to get <div>.

What to watch for


The details that decide whether a conversion is correct, and where information can be lost without any error being raised.

Named entities
The structural five (&amp; &lt; &gt; &quot; &apos;) are supported along with around eighty common named entities, &nbsp;, &copy;, &mdash;, &hellip;, typographic quotes, Greek letters, arrows and mathematical symbols. The full HTML5 set runs to well over two thousand names; anything unrecognised is left untouched and reported rather than silently deleted.
Numeric references always work
Decimal references like &#8212; and hexadecimal ones like &#x2014; are both handled, and cover every Unicode code point. If a named entity is not recognised, its numeric equivalent will be. Invalid code points and lone surrogates are rejected rather than producing corrupt output.
Multiple rounds for double-escaped text
Text escaped twice shows as &amp;lt;, decoding once gives &lt;, and decoding again gives <. Run the tool repeatedly until the output stops changing. Double escaping usually means two layers of code are both escaping the same value.
Why this does not use the browser DOM
Decoding via innerHTML would parse the input as HTML, which executes event handlers and inline scripts. That makes the obvious implementation a security hole. This tool maps entities directly, so hostile input is only ever data.

Limitations


  • Supports the common named entities rather than the complete HTML5 set of over 2,000; numeric references always work.
  • Double-escaped input needs more than one pass.
  • Processing happens in your browser, so very large inputs are bounded by available memory. Files above roughly 10 MB are handled but will feel slower, and multi-hundred-megabyte files are better suited to a command-line tool.

Questions


Why was an entity left unchanged?
It is not in the supported set. Rather than deleting text it cannot interpret, the tool leaves it and lists it in a warning. Numeric references always decode, so convert the entity to its code point if you need it.
Why does the output still contain &lt;?
The input was escaped more than once. Each pass removes one layer, so run it again.
Is decoding untrusted HTML safe here?
Yes. The input is never parsed as markup or inserted into the page, entities are mapped directly to characters, so nothing can execute.