Skip to content
Novus Examples
xml344 B

Sample XML

A well-formed XML catalogue with nested elements and attributes — for testing XML parsers and XPath queries.

Preview — first 14 linesxml
<?xml version="1.0" encoding="UTF-8"?>
<catalog>
  <product id="1">
    <name>Widget</name>
    <price currency="USD">9.99</price>
    <tags><tag>tools</tag><tag>hardware</tag></tags>
  </product>
  <product id="2">
    <name>Gadget</name>
    <price currency="USD">19.99</price>
    <tags><tag>electronics</tag></tags>
  </product>
</catalog>

Specifications

Elements
catalog › product › name/price/tags
Attributes
true
Declaration
UTF-8

What is a .xml file?

XML (Extensible Markup Language) is a verbose, self-describing markup language using nested tags, attributes, and namespaces to represent structured, hierarchical data. It supports schemas, entities, and validation and underlies many document and data formats. It remains common in enterprise, publishing, and interchange contexts.

How to use this file

Use an example XML file to test parsers, namespace and schema validation, XPath queries, and protection against entity-expansion and external-entity attacks.

How to use this file for testing

“Sample XML” is a deterministic Novus Examples fixture for Data import. Realistic faker-generated datasets with documented schemas for testing import and ETL flows.

Documented properties for this file: XML · 344 bytes. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

Data fixtures document their exact quirks — delimiters, encodings, null handling, schema, and row counts — in the spec table. Point your parser or importer at the file and assert it handles the documented edge cases; clean and deliberately-messy siblings make before/after diffs straightforward.

Code examples

import xml.etree.ElementTree as ET

tree = ET.parse("sample.xml")
root = tree.getroot()
print(root.tag, [c.tag for c in root][:5])

Generated by generation/data_structured.py. Free for any use, no attribution required — license.