Text Embeddings — 8-dim (NPY)
NumPy matrix twin of the Wave F 8-dim embeddings.
| dim_00 | dim_01 | dim_02 | dim_03 | dim_04 |
|---|---|---|---|---|
| -0.24161726236343384 | 0.2968849539756775 | -0.08661457151174545 | 0.2776573896408081 | 0.5728073716163635 |
| 0.38322508335113525 | -0.5825456976890564 | 0.3423929214477539 | 0.25020214915275574 | -0.3601084053516388 |
| 0.45138081908226013 | 0.10688069462776184 | 0.37255197763442993 | 0.5447877645492554 | -0.14542654156684875 |
| 0.11830972135066986 | 0.5287216901779175 | -0.2304946482181549 | -0.0516061894595623 | -0.6382369995117188 |
| 0.40608295798301697 | -0.3822888731956482 | 0.30053913593292236 | 0.5241363048553467 | 0.4103422462940216 |
| 0.3662438988685608 | -0.18557460606098175 | -0.10217946022748947 | -0.14612257480621338 | -0.7031015157699585 |
| -0.07396814972162247 | -0.16625788807868958 | 0.3832027316093445 | 0.0039348178543150425 | 0.1684892773628235 |
| -0.08848331868648529 | 0.539888322353363 | -0.6576524972915649 | -0.08510672301054001 | -0.2810226082801819 |
Specifications
- Shape
- 12x8
- Dtype
- float32
- Seed
- 314159
What is a .npy file?
NPY is NumPy's native binary format for a single array. A short header records the dtype, shape, and memory order, followed by the raw array bytes, so an array round-trips exactly without any text parsing. It is the standard way to persist embeddings, tensors, and numeric matrices in the Python data stack.
How to use this file
Use an example .npy to test array loaders (numpy.load), tensor and embedding pipelines, and converters between .npy, JSON, and columnar formats like Parquet.
How to use this file for testing
“Text Embeddings — 8-dim (NPY)” is a deterministic Novus Examples fixture for ML training data. Labelled, synthetic datasets in the shapes ML pipelines expect — JSONL for text tasks, image annotations, embeddings, and sample weights — for testing data loaders, tokenizers, and training tooling.
Documented properties for this file: seed 314159. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
AI/ML fixtures are fully synthetic with documented schemas — no real people or data. Test data loaders, tokenizers, annotation converters, embedding/vector stores, or eval-metric parsers against the known structure and fixed seeds.
These are labelled, training-shaped fixtures with a documented schema. Test your data loader, tokenizer, or format converter against it; every label and value is synthetic.
Related files
- jsonMultilingual Embeddings (JSON)EN/ES sentence embeddings for cross-lingual similarity tests.

- jsonQuantized int8 Embeddings (JSON)int8-quantised embedding vector for quantised vector search tests.

- jsonText Embeddings — 16-dim (JSON)A set of 24 L2-normalised 16-dimensional text embeddings as JSON — each record pairs an id and its source text with a float vector. A fixture for testing vector stores, similarity search, and embedding loaders. Parquet and .npy twins included.

- parquetText Embeddings — 16-dim (Parquet)The same 16-dimensional embeddings as Apache Parquet — id and text columns plus one column per dimension. The columnar twin, for testing analytics engines and Parquet-based vector pipelines.

- npyText Embeddings — 24×16 matrix (NumPy .npy)The embeddings as a raw NumPy array — a 24×16 float32 matrix in .npy format, loadable with numpy.load. The binary twin of the JSON and Parquet files, for testing tensor and matrix loaders.

- jsonlASR Digit Utterances Dataset (JSONL)JSON Lines ASR training/eval set for the Wave B synthetic digit utterances — each row points at a clean WAV and carries the expected transcript.

Generated by generation/ai_wave_f.py. Free for any use, no attribution required — license.