Sample XML
A well-formed XML catalogue with nested elements and attributes — for testing XML parsers and XPath queries.
<?xml version="1.0" encoding="UTF-8"?>
<catalog>
<product id="1">
<name>Widget</name>
<price currency="USD">9.99</price>
<tags><tag>tools</tag><tag>hardware</tag></tags>
</product>
<product id="2">
<name>Gadget</name>
<price currency="USD">19.99</price>
<tags><tag>electronics</tag></tags>
</product>
</catalog>
Specifications
- Elements
- catalog › product › name/price/tags
- Attributes
- true
- Declaration
- UTF-8
What is a .xml file?
XML (Extensible Markup Language) is a verbose, self-describing markup language using nested tags, attributes, and namespaces to represent structured, hierarchical data. It supports schemas, entities, and validation and underlies many document and data formats. It remains common in enterprise, publishing, and interchange contexts.
How to use this file
Use an example XML file to test parsers, namespace and schema validation, XPath queries, and protection against entity-expansion and external-entity attacks.
How to use this file for testing
“Sample XML” is a deterministic Novus Examples fixture for Data import. Realistic faker-generated datasets with documented schemas for testing import and ETL flows.
Documented properties for this file: XML · 344 bytes. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Data fixtures document their exact quirks — delimiters, encodings, null handling, schema, and row counts — in the spec table. Point your parser or importer at the file and assert it handles the documented edge cases; clean and deliberately-messy siblings make before/after diffs straightforward.
Code examples
import xml.etree.ElementTree as ET
tree = ET.parse("sample.xml")
root = tree.getroot()
print(root.tag, [c.tag for c in root][:5])Related files
- csv10,000-Row CSVA CSV with 10,000 data rows — for testing streaming parsers, memory handling, and import performance.

- logApache Access Log — Combined Format (log)An Apache access log in the Combined Log Format — the Common fields plus referer and user-agent. A realistic web-server log for testing access-log parsers, analytics, and grok patterns.

- logApache Access Log — Common Format (log)An Apache access log in the Common Log Format (CLF) — client, timestamp, request line, status, and byte count. Deterministic, with reserved documentation IPs. Paired with a Combined-format twin for testing log parsers.

- csvBank Transactions (CSV, 60 rows)A bank-transaction statement — 60 debits and credits across three accounts (masked numbers) with running balances, categories, and merchants. Synthetic data for testing statement parsers, categorisation, and reconciliation.

- jsonBank Transactions (JSON, 60 records)The bank transactions as a JSON array — the format twin of the CSV, for import and reconciliation testing.

- csvClean CSVA clean, well-formed CSV with a header and 20 rows — the baseline case for CSV parser testing.

Generated by generation/data_structured.py. Free for any use, no attribution required — license.