- How the row element is chosen
- The document is walked looking for the first element that appears more than once at the same level with object-like content. That is nearly always the intended row. If your document nests several repeated elements, the outermost is used, extract the inner fragment yourself if you need a different one.
- Nested elements become dotted columns
- Child elements inside a row are flattened into dotted names, so an <author><name> inside <book> becomes the column author.name. Attributes appear as columns prefixed with @_, keeping them distinct from child elements of the same name.
- Rows with different children still line up
- The header is the union of every field seen across all rows, so a row missing an element simply gets an empty cell. No row is ever shorter than the header, which keeps the file valid for strict parsers.
- What cannot be represented
- Mixed content, comments and namespace semantics are all lost. A repeated child element inside a single row has no column form and is serialised as text. Types disappear entirely, since both XML content and CSV cells are untyped text.