Skip to content
Novus Examples
json659 B

Multilingual Embeddings (JSON)

EN/ES sentence embeddings for cross-lingual similarity tests.

Preview — first 33 linesjson
[
  {
    "id": "t00",
    "text": "Hello sample",
    "lang": "en",
    "embedding": [
      -0.24161726236343384,
      0.2968849539756775,
      -0.08661457151174545,
      0.2776573896408081,
      0.5728073716163635,
      0.042604897171258926,
      -0.3045431971549988,
      -0.5884001851081848
    ]
  },
  {
    "id": "t01",
    "text": "Hola muestra",
    "lang": "es",
    "embedding": [
      0.38322508335113525,
      -0.5825456976890564,
      0.3423929214477539,
      0.25020214915275574,
      -0.3601084053516388,
      -0.08197279274463654,
      -0.4132791757583618,
      0.16354766488075256
    ]
  }
]

Specifications

Languages
en, es
Dimensions
8

What is a .json file?

JSON (JavaScript Object Notation) is a lightweight, text-based data-interchange format representing objects, arrays, strings, numbers, booleans, and null. It is language-independent, human-readable, and the dominant format for web APIs and configuration. It requires a single well-formed root value.

How to use this file

Use an example JSON file to test parsers and serializers, schema validation, Unicode and number-precision handling, and API request or response processing.

How to use this file for testing

“Multilingual Embeddings (JSON)” is a deterministic Novus Examples fixture for ML training data. Labelled, synthetic datasets in the shapes ML pipelines expect — JSONL for text tasks, image annotations, embeddings, and sample weights — for testing data loaders, tokenizers, and training tooling.

Documented properties for this file: JSON · 659 bytes. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

AI/ML fixtures are fully synthetic with documented schemas — no real people or data. Test data loaders, tokenizers, annotation converters, embedding/vector stores, or eval-metric parsers against the known structure and fixed seeds.

These are labelled, training-shaped fixtures with a documented schema. Test your data loader, tokenizer, or format converter against it; every label and value is synthetic.

Code examples

import json

with open("multilingual-embeddings.json") as f:
    data = json.load(f)
print(type(data), len(data))

Generated by generation/ai_wave_f.py. Free for any use, no attribution required — license.