Convert v2 ORC Employee Table Source
Binary orc source for the five-row P8 employee conversion table, preserving ids, names, departments, booleans, and scores. Stable P8 artifact p8-convert-orc-source.
| id | name | department | active | score | joined | |
|---|---|---|---|---|---|---|
| 1001 | Ada Lovelace | ada.lovelace@example.com | Engineering | true | 98.5 | 2021-03-01 |
| 1002 | Alan Turing | alan.turing@example.com | Research | true | 95 | 2020-06-15 |
| 1003 | Grace Hopper | grace.hopper@example.com | Engineering | false | 91.2 | 2019-11-20 |
| 1004 | Katherine Johnson | katherine.johnson@example.com | Operations | true | 96.8 | 2022-01-10 |
| 1005 | Edsger Dijkstra | edsger.dijkstra@example.com | Research | false | 89.4 | 2018-09-05 |
Specifications
- Rows
- 5
- Columns
- 7
- Source Format
- orc
- Delivery Mode
- download-only
- Provider
- converter-v2
- Provenance
- Synthetic deterministic P8 fixture generated by generation/p8_content.py; seed namespace 2026082300
- Fixture Reserve
- convert-v2
Testing contract
Expected to pass- Scenario
- Read the ORC artifact and project the five contract columns in row order.
- Expected result
- Five rows and all seven columns match ids 1001 through 1005; Ada's email and 2021-03-01 join date survive, Grace is inactive, and scores remain exact.
What is a .orc file?
Apache ORC (Optimized Row Columnar, .orc) is a binary columnar format from the Hadoop ecosystem. It stores data in stripes with lightweight indexes, per-column compression, and embedded statistics, and is common in Hive and big-data pipelines.
How to use this file
Use an example .orc file to test ORC readers, stripe and index handling, and ORC-to-Parquet/CSV conversion.
How to use this file for testing
“Convert v2 ORC Employee Table Source” is a deterministic Novus Examples fixture for Conversion testing, Data engineering, Serialization testing. The same content exported across many formats and linked as a group, so you can convert one and diff against the expected twin.
Documented properties for this file: 5 rows · 7 columns. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Data fixtures document their exact quirks — delimiters, encodings, null handling, schema, and row counts — in the spec table. Point your parser or importer at the file and assert it handles the documented edge cases; clean and deliberately-messy siblings make before/after diffs straightforward.
Related files
- pbConvert v2 Protobuf Creator RecordValid Protobuf wire record with string, uint32, and string fields for schema-guided decoding. Stable P8 artifact p8-convert-protobuf-source.

- jsonConvert v2 Protobuf Expected JSONExpected JSON semantic result for the schema-guided Protobuf decode. Stable P8 artifact p8-convert-protobuf-expected.

- pbConvert v2 Protobuf Missing-Schema Controlled FailureValid wire bytes intentionally supplied without a matching schema so tools must preserve unknown field 15. Stable P8 artifact p8-convert-protobuf-missing-schema.

- protoConvert v2 Protobuf Schema CompanionProto3 schema companion declaring the three fields used by the binary creator record. Stable P8 artifact p8-convert-protobuf-schema.

- avroAvro — Row Binary + SchemaThe same records as Apache Avro — a compact row-based binary format that embeds its own schema, widely used in Kafka pipelines. For testing Avro decoders and schema evolution.

- bsonBSON — Binary JSON (MongoDB)The records as BSON — the binary-JSON encoding MongoDB stores documents in. For testing BSON decoders and JSON↔BSON conversion.

Generated by generation/p8_content.py. Free for any use, no attribution required — license.