HDF5 — Hierarchical Scientific Data
A small HDF5 file with a compound 'employees' dataset and a numeric 'readings' grid — the hierarchical format used across science and ML. For testing h5py/HDF5 readers and conversion.
| idint64 | namestring | emailstring | departmentstring | activebool | scoredouble | joineddate |
|---|---|---|---|---|---|---|
| 1001 | Ada Lovelace | ada.lovelace@example.com | Engineering | true | 98.5 | 2021-03-01 |
| 1002 | Alan Turing | alan.turing@example.com | Research | true | 95 | 2020-06-15 |
| 1003 | Grace Hopper | grace.hopper@example.com | Engineering | false | 91.2 | 2019-11-20 |
| 1004 | Katherine Johnson | katherine.johnson@example.com | Operations | true | 96.8 | 2022-01-10 |
| 1005 | Edsger Dijkstra | edsger.dijkstra@example.com | Research | false | 89.4 | 2018-09-05 |
Specifications
- Rows
- 5
- Columns
- 7
- Format
- HDF5
- Datasets
- 2
- Groups
- 1
What is a .h5 file?
HDF5 (.h5) is a binary container format for large, heterogeneous scientific data. It stores multidimensional arrays (datasets) in a hierarchical group structure with attributes and chunked, compressed storage, and is standard in ML, physics, and geoscience.
How to use this file
Use an example .h5 file to test HDF5 readers (h5py, PyTables), group and dataset traversal, and attribute extraction.
How to use this file for testing
“HDF5 — Hierarchical Scientific Data” is a deterministic Novus Examples fixture for Conversion testing, Data engineering. The same content exported across many formats and linked as a group, so you can convert one and diff against the expected twin.
Documented properties for this file: 5 rows · 7 columns · HDF5. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Data fixtures document their exact quirks — delimiters, encodings, null handling, schema, and row counts — in the spec table. Point your parser or importer at the file and assert it handles the documented edge cases; clean and deliberately-messy siblings make before/after diffs straightforward.
Related files
- avroAvro — Row Binary + SchemaThe same records as Apache Avro — a compact row-based binary format that embeds its own schema, widely used in Kafka pipelines. For testing Avro decoders and schema evolution.

- featherFeather — Arrow IPC TableThe same table as Feather (Arrow IPC file) — the zero-copy on-disk form of an Apache Arrow table. For testing Arrow readers and fast columnar interchange.

- orcORC — Columnar TableThe same employee table as Apache ORC — the columnar format common in the Hive/Hadoop ecosystem. For testing ORC readers and Parquet↔ORC conversion.

- parquetParquet — Columnar TableA small employee table as Apache Parquet — the columnar format at the heart of modern data lakes and analytics. For testing Parquet readers (pandas, Spark, DuckDB) and conversion.

- orcConvert v2 ORC Employee Table SourceBinary orc source for the five-row P8 employee conversion table, preserving ids, names, departments, booleans, and scores. Stable P8 artifact p8-convert-orc-source.

- csvConvert v2 ORC Expected CSVCsv semantic reference for the five-row P8 employee conversion table, preserving ids, names, departments, booleans, and scores. Stable P8 artifact p8-convert-orc-csv.

Generated by generation/data_binary.py. Free for any use, no attribution required — license.