- Attributes get an @_ prefix
- JSON has no attribute concept, so <book id="1"> becomes {"book": {"@_id": "1"}}. The prefix prevents a collision when an element also has a child named id, without it, one would overwrite the other. Attribute values are always kept as strings even when type inference is on, because attributes are overwhelmingly identifiers where turning "007" into 7 would corrupt the value.
- One element or an array? It depends on the document
- This is the single biggest gotcha. <items><i>x</i></items> yields a string, while <items><i>x</i><i>y</i></items> yields an array. Nothing in the XML itself says whether a list was intended, so code consuming the JSON must handle both, normalise with Array.isArray(v) ? v : [v] before iterating. An XSD would resolve the ambiguity, but a converter working from the instance document alone cannot.
- Namespaces stay in the key name
- A prefixed element like <soap:Body> becomes the key "soap:Body", and xmlns declarations become ordinary @_xmlns attributes. JSON has no namespace mechanism, so the prefix is preserved as literal text. Two elements in different namespaces that share a local name remain distinct only because their prefixes differ, if the document uses different prefixes for the same namespace, the JSON will not reflect that they are equivalent.
- Mixed content appears under #text
- When an element contains both text and children, as in <p>Hello <b>world</b></p>, the loose text is stored under the "#text" key alongside the child keys. Note that the original interleaving is lost: JSON object keys cannot record that the text came before the <b> element, so mixed-content documents like prose markup do not round-trip faithfully.
- CDATA, comments and entities
- CDATA sections become plain strings, the wrapper exists only to escape markup, so it carries no data. XML comments are dropped, as JSON has no equivalent. The five predefined entities (< > & " ') are resolved to their characters. External entity references are never resolved, which blocks the XXE class of attack, and entity-expansion bombs are rejected rather than allowed to exhaust memory.
- Empty elements
- A self-closing <tag/> and an empty <tag></tag> both become an empty string, since XML treats them identically. There is no way to express an XML-level "nil" in JSON other than by convention, so a null in your data usually needs xsi:nil handling that this converter leaves as a plain attribute.