HTML to Markdown

Turn HTML into clean, readable Markdown.

Reverse
Input
HTML input
Output
Result
Options

About this tool


Converting HTML to Markdown is lossy by nature, because HTML can express far more than Markdown can. The goal is not a perfect representation but readable Markdown that keeps the content and its meaningful structure.

This is what you want for turning a web page into documentation, cleaning up content pasted out of a rich text editor, or getting rendered HTML back into a source format you can edit by hand.

How to use it

  1. Paste or upload your htmlDrop a file onto the input pane, use the file picker, or paste the text directly.
  2. Adjust the options if neededThe defaults suit most input; open Options to change the behaviour.
  3. Convert to MarkdownPress Convert to Markdown, or use Ctrl+Enter (Cmd+Enter on macOS).
  4. Copy or downloadCopy the result, or download it as a .md file.

Worked examples


Each example below is executed against this tool by the test suite, so what you see is what the tool actually produces.

A content fragment

Input

<h2>Title</h2><p>Some <strong>bold</strong> text.</p><ul><li>one</li><li>two</li></ul>

Output

## Title

Some **bold** text.

- one
- two

Semantic elements map to Markdown; the list becomes hyphen bullets.

What to watch for


The details that decide whether a conversion is correct, and where information can be lost without any error being raised.

Non-content elements are discarded
Script, style and noscript blocks are removed entirely, along with comments. Structural wrappers like div, section and span have no Markdown equivalent and are unwrapped, leaving their contents. The result is content rather than markup.
What converts cleanly
Headings become # levels, strong and em become ** and *, code and pre become backticks and fences, links and images become bracket syntax, lists become hyphens or numbers, and blockquotes become >. A language class on a code block is preserved as the fence hint.
Tables convert only if they are simple
A regular grid becomes a pipe table with aligned columns. Markdown tables cannot express merged cells, so colspan and rowspan are lost and those tables come out misaligned. Tables used purely for layout, common in email HTML, produce poor results because they were never tabular data.
Styling and attributes disappear
Markdown has no way to express inline styles, classes, ids, colours, or font choices, so all are dropped. Content that relied on styling for meaning (colour-coded text, for instance) loses that meaning. Only semantic markup survives.
Round-tripping is not exact
Converting HTML to Markdown and back gives you equivalent content, not the original HTML. Attributes, wrapper elements and whitespace differ. Treat whichever format you edit as the source of truth.

Limitations


  • Inline styles, classes, ids and colours are discarded.
  • Tables with merged cells cannot be represented.
  • Nested structures deeper than a few levels may flatten.
  • Processing happens in your browser, so very large inputs are bounded by available memory. Files above roughly 10 MB are handled but will feel slower, and multi-hundred-megabyte files are better suited to a command-line tool.

Questions


Why is my styling gone?
Markdown cannot express inline styles, classes or colours, so they are dropped. Only semantic structure survives the conversion.
Why is my table misaligned?
It probably uses merged cells. Markdown tables have no colspan or rowspan, so such tables cannot be represented faithfully.
Can I get my exact HTML back?
No. The round trip preserves content, not markup. Attributes and wrapper elements are lost, so pick one format as your source.