Skip to content
Novus Examples
parquet31.5 KB

E-commerce Orders (Parquet, 2000 rows)

The e-commerce orders table as Apache Parquet — the columnar twin, for testing analytics engines (pandas, DuckDB, Spark).

Preview — schema + first 8 rowsparquet
order_idcustomer_idproduct_idquantitytotalstatusorder_date
11501712705.98paid2025-10-19
217436125.59refunded2025-06-09
33112444.46shipped2025-04-27
4293514881.44delivered2025-08-22
543112731105.89shipped2025-03-12
62951905234.85delivered2025-05-03
73365252229.65refunded2025-05-26
8236544207pending2025-04-29
Decoded Parquet — first 8 of 2000 rows.

Specifications

Rows
2000
Columns
7
Format
Apache Parquet
Domain
e-commerce

What is a .parquet file?

Apache Parquet (.parquet) is a binary, columnar storage format for analytical data. It stores each column separately with per-column compression and encoding, embeds a schema and statistics, and is the de-facto standard for data lakes and engines like Spark, DuckDB, and pandas/pyarrow.

How to use this file

Use an example .parquet file to test columnar readers (pyarrow, DuckDB, Spark), schema and predicate-pushdown handling, and Parquet-to-CSV/JSON converters.

How to use this file for testing

“E-commerce Orders (Parquet, 2000 rows)” is a deterministic Novus Examples fixture for Data engineering, Conversion testing. Columnar (Parquet/ORC/Feather), row (Avro), and messaging (MessagePack/CBOR/Protobuf) formats plus star-schema and log data — for testing ETL, data-lake ingestion, and warehouse loaders.

Documented properties for this file: 2,000 rows · 7 columns · Apache Parquet. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

Data fixtures document their exact quirks — delimiters, encodings, null handling, schema, and row counts — in the spec table. Point your parser or importer at the file and assert it handles the documented edge cases; clean and deliberately-messy siblings make before/after diffs straightforward.

Code examples

import pandas as pd  # pip install pyarrow

df = pd.read_parquet("orders.parquet")
print(df.head())
print(df.dtypes)

Generated by generation/data_realworld.py. Free for any use, no attribution required — license.