XML to JSON

Turn XML into JSON, with attributes and text nodes preserved.

Reverse
Input
XML input
Output
Result
Options

Attributes are prefixed with @_ so they stay distinct from elements.

Values with leading zeros stay strings so IDs are preserved.

Spaces per level of nesting.

About this tool


XML and JSON have genuinely different data models, so this conversion always involves choices. XML distinguishes attributes from child elements, allows text and elements to be mixed inside the same tag, and has namespaces. None of which JSON has. This converter makes each choice explicit rather than pretending the mapping is obvious.

The most important consequence to understand before you rely on the output: XML does not distinguish "one item" from "a list of one item". An element appearing once becomes a single value, and the same element repeated becomes an array, so two documents following the same schema can produce different JSON shapes.

How to use it

  1. Paste or upload XMLDrop an .xml file or paste the markup directly.
  2. Decide about attributesKeep them (prefixed with @_ so they stay distinct from elements) or discard them entirely.
  3. ConvertPress Convert to JSON, or use Ctrl+Enter.
  4. Read the warningsThey cover attribute prefixes, namespaces, mixed content and the single-vs-array behaviour.
  5. Copy or downloadSave the JSON or copy it to your clipboard.

Worked examples


Each example below is executed against this tool by the test suite, so what you see is what the tool actually produces.

Attributes and repeated elements

Input

<catalog>
  <book id="1" lang="en">
    <title>XML Basics</title>
  </book>
  <book id="2" lang="fr">
    <title>XML Avancé</title>
  </book>
</catalog>

Output

{
  "catalog": {
    "book": [
      {
        "title": "XML Basics",
        "@_id": "1",
        "@_lang": "en"
      },
      {
        "title": "XML Avancé",
        "@_id": "2",
        "@_lang": "fr"
      }
    ]
  }
}

Two <book> elements become an array. With only one book, "book" would be a single object instead.

Mixed content

Input

<p>Hello <b>world</b> and goodbye</p>

Output

{
  "p": {
    "b": "world",
    "#text": "Hello and goodbye"
  }
}

The loose text lands under #text, but its original position around <b> is not recorded.

What to watch for


The details that decide whether a conversion is correct, and where information can be lost without any error being raised.

Attributes get an @_ prefix
JSON has no attribute concept, so <book id="1"> becomes {"book": {"@_id": "1"}}. The prefix prevents a collision when an element also has a child named id, without it, one would overwrite the other. Attribute values are always kept as strings even when type inference is on, because attributes are overwhelmingly identifiers where turning "007" into 7 would corrupt the value.
One element or an array? It depends on the document
This is the single biggest gotcha. <items><i>x</i></items> yields a string, while <items><i>x</i><i>y</i></items> yields an array. Nothing in the XML itself says whether a list was intended, so code consuming the JSON must handle both, normalise with Array.isArray(v) ? v : [v] before iterating. An XSD would resolve the ambiguity, but a converter working from the instance document alone cannot.
Namespaces stay in the key name
A prefixed element like <soap:Body> becomes the key "soap:Body", and xmlns declarations become ordinary @_xmlns attributes. JSON has no namespace mechanism, so the prefix is preserved as literal text. Two elements in different namespaces that share a local name remain distinct only because their prefixes differ, if the document uses different prefixes for the same namespace, the JSON will not reflect that they are equivalent.
Mixed content appears under #text
When an element contains both text and children, as in <p>Hello <b>world</b></p>, the loose text is stored under the "#text" key alongside the child keys. Note that the original interleaving is lost: JSON object keys cannot record that the text came before the <b> element, so mixed-content documents like prose markup do not round-trip faithfully.
CDATA, comments and entities
CDATA sections become plain strings, the wrapper exists only to escape markup, so it carries no data. XML comments are dropped, as JSON has no equivalent. The five predefined entities (&lt; &gt; &amp; &quot; &apos;) are resolved to their characters. External entity references are never resolved, which blocks the XXE class of attack, and entity-expansion bombs are rejected rather than allowed to exhaust memory.
Empty elements
A self-closing <tag/> and an empty <tag></tag> both become an empty string, since XML treats them identically. There is no way to express an XML-level "nil" in JSON other than by convention, so a null in your data usually needs xsi:nil handling that this converter leaves as a plain attribute.

Limitations


  • A single element and a one-element list are indistinguishable, so output shape varies with the data.
  • Mixed-content ordering is lost; text and child elements cannot be interleaved in JSON.
  • Comments, DTDs and processing instructions are discarded.
  • Namespace prefixes are preserved as literal text, with no namespace resolution.
  • Processing happens in your browser, so very large inputs are bounded by available memory. Files above roughly 10 MB are handled but will feel slower, and multi-hundred-megabyte files are better suited to a command-line tool.

Questions


Why does the same schema give me different JSON shapes?
Because XML cannot distinguish a single item from a one-item list. One <item> produces a value; two produce an array. Normalise in your code with Array.isArray(v) ? v : [v].
Can I get rid of the @_ prefixes?
Turn off "Keep attributes" to discard attributes entirely. The prefix itself is not configurable, because it exists to prevent silent collisions between an attribute and a child element of the same name.
Is this safe against XXE and entity bombs?
Yes. External entity references are never resolved, so a SYSTEM entity pointing at a local file returns nothing. Recursive entity expansion is rejected rather than allowed to consume memory. Everything runs in your browser, so there is no server-side file system to reach in any case.
Will converting back to XML give me the original document?
Usually equivalent, rarely identical. Comments, the DTD, processing instructions and CDATA wrappers are not preserved, and mixed-content ordering is lost.