Skip to content
Novus Examples
json155 B

BM25 Retrieval Scores (JSON)

BM25 score snapshot for comparing neural rerankers against a lexical baseline.

Preview, first 14 linesjson
{
  "query": "sample",
  "scores": [
    {
      "doc": "d1",
      "bm25": 2.4
    },
    {
      "doc": "d2",
      "bm25": 1.1
    }
  ]
}

Specifications

Task
retrieval baseline

Testing contract

Expected to pass
Scenario
Exercise BM25 Retrieval Scores (JSON) in its ranking workflow. BM25 score snapshot for comparing neural rerankers against a lexical baseline.
Expected result
top-level keys are query, scores; array lengths: scores=2. Declared feature checks: task=retrieval baseline.

What is a .json file?

JSON (JavaScript Object Notation) is a lightweight, text-based data-interchange format representing objects, arrays, strings, numbers, booleans, and null. It is language-independent, human-readable, and the dominant format for web APIs and configuration. It requires a single well-formed root value.

How to use this file

Use an example JSON file to test parsers and serializers, schema validation, Unicode and number-precision handling, and API request or response processing.

How to use this file for testing

“BM25 Retrieval Scores (JSON)” is a deterministic Novus Examples fixture for Model evaluation, JSON parsing. Benchmark results, confusion matrices, ROC curves, and classification reports in CSV and JSON, for testing eval dashboards, metric parsers, and leaderboard importers.

Documented properties for this file: task: retrieval baseline. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such, expect parsers to fail loudly rather than silently accept them.

AI/ML fixtures are fully synthetic with documented schemas, no real people or data. Test data loaders, tokenizers, annotation converters, embedding/vector stores, or eval-metric parsers against the known structure and fixed seeds.

Code examples

import json

with open("bm25-scores.json") as f:
    data = json.load(f)
print(type(data), len(data))

Generated by generation/ai_wave_f.py. Free for any use, no attribution required, license.