Parquet — Null-heavy Columns
Tiny Parquet table with null name/score/note cells — tests null handling in columnar readers.
| id | name | score | note |
|---|---|---|---|
| 1 | Alpha | 90 | null |
| 2 | null | null | missing name |
| 3 | Gamma | 88.5 | null |
Specifications
- Edge
- null cells
- Rows
- 3
- Seed
- 42042
What is a .parquet file?
Apache Parquet (.parquet) is a binary, columnar storage format for analytical data. It stores each column separately with per-column compression and encoding, embeds a schema and statistics, and is the de-facto standard for data lakes and engines like Spark, DuckDB, and pandas/pyarrow.
How to use this file
Use an example .parquet file to test columnar readers (pyarrow, DuckDB, Spark), schema and predicate-pushdown handling, and Parquet-to-CSV/JSON converters.
How to use this file for testing
“Parquet — Null-heavy Columns” is a deterministic Novus Examples fixture for Data import, Conversion testing. Realistic faker-generated datasets with documented schemas for testing import and ETL flows.
Documented properties for this file: seed 42042 · 3 rows. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Data fixtures document their exact quirks — delimiters, encodings, null handling, schema, and row counts — in the spec table. Point your parser or importer at the file and assert it handles the documented edge cases; clean and deliberately-messy siblings make before/after diffs straightforward.
Code examples
import pandas as pd # pip install pyarrow
df = pd.read_parquet("nulls-heavy.parquet")
print(df.head())
print(df.dtypes)Related files
- featherFeather — Dictionary-encoded ColumnFeather file with dictionary-encoded strings — Arrow IPC edge-case fixture.

- parquetParquet — Decimal as String ColumnAmounts stored as strings in Parquet — common ingestion edge case for ETL parsers.

- parquetParquet — Dictionary-encoded ColumnParquet with dictionary-encoded string column — tests dictionary page decoding.

- parquetParquet — Duplicate Key ColumnParquet with repeated key values — tests join/aggregation edge cases.

- parquetParquet — Empty TableEmpty Parquet file with schema but no rows — edge case for readers.

- parquetParquet — List ColumnParquet with a list-of-string column — tests repeated-field decoding.

Generated by generation/data_wave_f.py. Free for any use, no attribution required — license.