Skip to content
Novus Examples
parquet765 B

Parquet — Decimal as String Column

Amounts stored as strings in Parquet — common ingestion edge case for ETL parsers.

Preview — schema + first 2 rowsparquet
idnamescorenote
1nullnullnull
2nullnullnull
First rows with null cells.

Specifications

Edge
decimal stored as string
Rows
2
Seed
42042

What is a .parquet file?

Apache Parquet (.parquet) is a binary, columnar storage format for analytical data. It stores each column separately with per-column compression and encoding, embeds a schema and statistics, and is the de-facto standard for data lakes and engines like Spark, DuckDB, and pandas/pyarrow.

How to use this file

Use an example .parquet file to test columnar readers (pyarrow, DuckDB, Spark), schema and predicate-pushdown handling, and Parquet-to-CSV/JSON converters.

How to use this file for testing

“Parquet — Decimal as String Column” is a deterministic Novus Examples fixture for Data import, Conversion testing. Realistic faker-generated datasets with documented schemas for testing import and ETL flows.

Documented properties for this file: seed 42042 · 2 rows. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

Data fixtures document their exact quirks — delimiters, encodings, null handling, schema, and row counts — in the spec table. Point your parser or importer at the file and assert it handles the documented edge cases; clean and deliberately-messy siblings make before/after diffs straightforward.

Code examples

import pandas as pd  # pip install pyarrow

df = pd.read_parquet("decimal-as-string.parquet")
print(df.head())
print(df.dtypes)

Generated by generation/data_wave_f.py. Free for any use, no attribution required — license.