Parquet — Dictionary-encoded Column
Parquet with dictionary-encoded string column — tests dictionary page decoding.
| id | name | score | note |
|---|---|---|---|
| null | null | null | null |
| null | null | null | null |
| null | null | null | null |
| null | null | null | null |
| null | null | null | null |
Specifications
- Edge
- dictionary encoding
- Rows
- 5
- Seed
- 42042
What is a .parquet file?
Apache Parquet (.parquet) is a binary, columnar storage format for analytical data. It stores each column separately with per-column compression and encoding, embeds a schema and statistics, and is the de-facto standard for data lakes and engines like Spark, DuckDB, and pandas/pyarrow.
How to use this file
Use an example .parquet file to test columnar readers (pyarrow, DuckDB, Spark), schema and predicate-pushdown handling, and Parquet-to-CSV/JSON converters.
How to use this file for testing
“Parquet — Dictionary-encoded Column” is a deterministic Novus Examples fixture for Data import, Conversion testing. Realistic faker-generated datasets with documented schemas for testing import and ETL flows.
Documented properties for this file: seed 42042 · 5 rows. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Data fixtures document their exact quirks — delimiters, encodings, null handling, schema, and row counts — in the spec table. Point your parser or importer at the file and assert it handles the documented edge cases; clean and deliberately-messy siblings make before/after diffs straightforward.
Code examples
import pandas as pd # pip install pyarrow
df = pd.read_parquet("dictionary-encoded.parquet")
print(df.head())
print(df.dtypes)Related files
- jsonColumnar Nulls Schema (JSON)JSON description of nullable columns in the null-heavy columnar fixtures.

- featherFeather — Dictionary-encoded ColumnFeather file with dictionary-encoded strings — Arrow IPC edge-case fixture.

- featherFeather — Null-heavy ColumnsFeather/Arrow IPC twin of the null-heavy table — grouped with the Parquet nulls fixture.

- parquetParquet — Decimal as String ColumnAmounts stored as strings in Parquet — common ingestion edge case for ETL parsers.

- parquetParquet — Duplicate Key ColumnParquet with repeated key values — tests join/aggregation edge cases.

- parquetParquet — Empty TableEmpty Parquet file with schema but no rows — edge case for readers.

Generated by generation/data_wave_f.py. Free for any use, no attribution required — license.