Skip to content
Novus Examples
parquet765 B

Parquet — Decimal as String Column

Amounts stored as strings in Parquet — common ingestion edge case for ETL parsers.

Preview: schema + first 2 rowsparquet
idnamescorenote
1nullnullnull
2nullnullnull
First rows with null cells.

Specifications

Edge
decimal stored as string
Rows
2
Seed
42042

Testing contract

Expected to pass
Scenario
Exercise Parquet — Decimal as String Column in its columnar workflow. Amounts stored as strings in Parquet — common ingestion edge case for ETL parsers.
Expected result
2 rows, 2 columns; fields: id: int64; amount_str: string; column null counts=[0, 0]. Declared feature checks: edge=decimal stored as string.

What is a .parquet file?

Apache Parquet (.parquet) is a binary, columnar storage format for analytical data. It stores each column separately with per-column compression and encoding, embeds a schema and statistics, and is the de-facto standard for data lakes and engines like Spark, DuckDB, and pandas/pyarrow.

How to use this file

Use an example .parquet file to test columnar readers (pyarrow, DuckDB, Spark), schema and predicate-pushdown handling, and Parquet-to-CSV/JSON converters.

How to use this file for testing

“Parquet — Decimal as String Column” is a deterministic Novus Examples fixture for Data import, Conversion testing. Realistic faker-generated datasets with documented schemas for testing import and ETL flows.

Documented properties for this file: seed 42042 · 2 rows. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such, expect parsers to fail loudly rather than silently accept them.

Data fixtures document their exact quirks (delimiters, encodings, null handling, schema, and row counts) in the spec table. Point your parser or importer at the file and assert it handles the documented edge cases; clean and deliberately-messy siblings make before/after diffs straightforward.

Code examples

import pandas as pd  # pip install pyarrow

df = pd.read_parquet("decimal-as-string.parquet")
print(df.head())
print(df.dtypes)

Generated by generation/data_wave_f.py. Free for any use, no attribution required, license.