Multilingual Embeddings (JSON)
EN/ES sentence embeddings for cross-lingual similarity tests.
[
{
"id": "t00",
"text": "Hello sample",
"lang": "en",
"embedding": [
-0.24161726236343384,
0.2968849539756775,
-0.08661457151174545,
0.2776573896408081,
0.5728073716163635,
0.042604897171258926,
-0.3045431971549988,
-0.5884001851081848
]
},
{
"id": "t01",
"text": "Hola muestra",
"lang": "es",
"embedding": [
0.38322508335113525,
-0.5825456976890564,
0.3423929214477539,
0.25020214915275574,
-0.3601084053516388,
-0.08197279274463654,
-0.4132791757583618,
0.16354766488075256
]
}
]
Specifications
- Languages
- en, es
- Dimensions
- 8
What is a .json file?
JSON (JavaScript Object Notation) is a lightweight, text-based data-interchange format representing objects, arrays, strings, numbers, booleans, and null. It is language-independent, human-readable, and the dominant format for web APIs and configuration. It requires a single well-formed root value.
How to use this file
Use an example JSON file to test parsers and serializers, schema validation, Unicode and number-precision handling, and API request or response processing.
How to use this file for testing
“Multilingual Embeddings (JSON)” is a deterministic Novus Examples fixture for ML training data. Labelled, synthetic datasets in the shapes ML pipelines expect — JSONL for text tasks, image annotations, embeddings, and sample weights — for testing data loaders, tokenizers, and training tooling.
Documented properties for this file: JSON · 659 bytes. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
AI/ML fixtures are fully synthetic with documented schemas — no real people or data. Test data loaders, tokenizers, annotation converters, embedding/vector stores, or eval-metric parsers against the known structure and fixed seeds.
These are labelled, training-shaped fixtures with a documented schema. Test your data loader, tokenizer, or format converter against it; every label and value is synthetic.
Code examples
import json
with open("multilingual-embeddings.json") as f:
data = json.load(f)
print(type(data), len(data))Related files
- jsonEmbedding ID Mapping (JSON)ID list and model metadata for the Wave F embedding set.

- jsonQuantized int8 Embeddings (JSON)int8-quantised embedding vector for quantised vector search tests.

- jsonText Embeddings — 16-dim (JSON)A set of 24 L2-normalised 16-dimensional text embeddings as JSON — each record pairs an id and its source text with a float vector. A fixture for testing vector stores, similarity search, and embedding loaders. Parquet and .npy twins included.

- parquetText Embeddings — 16-dim (Parquet)The same 16-dimensional embeddings as Apache Parquet — id and text columns plus one column per dimension. The columnar twin, for testing analytics engines and Parquet-based vector pipelines.

- npyText Embeddings — 24×16 matrix (NumPy .npy)The embeddings as a raw NumPy array — a 24×16 float32 matrix in .npy format, loadable with numpy.load. The binary twin of the JSON and Parquet files, for testing tensor and matrix loaders.

- jsonText Embeddings — 8-dim (JSON)L2-normalised 8-dimensional text embeddings in JSON — for vector store loader tests.

Generated by generation/ai_wave_f.py. Free for any use, no attribution required — license.