Skip to content
Novus Examples
avro783 B

Avro — Row Binary + Schema

The same records as Apache Avro — a compact row-based binary format that embeds its own schema, widely used in Kafka pipelines. For testing Avro decoders and schema evolution.

Preview — schema + first 5 rowsavro
idint64namestringemailstringdepartmentstringactiveboolscoredoublejoineddate
1001Ada Lovelaceada.lovelace@example.comEngineeringtrue98.52021-03-01
1002Alan Turingalan.turing@example.comResearchtrue952020-06-15
1003Grace Hoppergrace.hopper@example.comEngineeringfalse91.22019-11-20
1004Katherine Johnsonkatherine.johnson@example.comOperationstrue96.82022-01-10
1005Edsger Dijkstraedsger.dijkstra@example.comResearchfalse89.42018-09-05
Decoded table — all 5 rows shown.

Specifications

Rows
5
Columns
7
Format
Apache Avro
Layout
row
Schema
embedded

What is a .avro file?

Apache Avro (.avro) is a compact binary, row-oriented data format that stores its JSON schema in the file header alongside the records, enabling schema evolution. It is widely used with Kafka and Hadoop for message and data serialization.

How to use this file

Use an example .avro file to test Avro readers, embedded-schema parsing, and schema-evolution or Avro-to-JSON tooling.

How to use this file for testing

“Avro — Row Binary + Schema” is a deterministic Novus Examples fixture for Conversion testing, Data engineering. The same content exported across many formats and linked as a group, so you can convert one and diff against the expected twin.

Documented properties for this file: 5 rows · 7 columns · schema: embedded. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

Data fixtures document their exact quirks — delimiters, encodings, null handling, schema, and row counts — in the spec table. Point your parser or importer at the file and assert it handles the documented edge cases; clean and deliberately-messy siblings make before/after diffs straightforward.

Code examples

import fastavro

with open("employees.avro", "rb") as f:
    for record in fastavro.reader(f):
        print(record)

Generated by generation/data_binary.py. Free for any use, no attribution required — license.