Model Benchmark Results (JSON)
A model-evaluation summary in JSON — per-task scores for a fictional model across sentiment, NER, summarization, translation, and QA, each with its metric and sample size. A fixture for testing eval dashboards and leaderboard importers.
{
"model": "novus-demo-1",
"created": "2026-01-01",
"suite": "mini-eval",
"results": [
{
"task": "sentiment",
"metric": "accuracy",
"score": 0.912,
"n": 500
},
{
"task": "ner",
"metric": "f1",
"score": 0.874,
"n": 300
},
{
"task": "summarization",
"metric": "rougeL",
"score": 0.381,
"n": 200
},
{
"task": "translation_en_es",
"metric": "bleu",
"score": 0.336,
"n": 400
},
{
"task": "qa_extractive",
"metric": "exact_match",
"score": 0.685,
"n": 350
}
]
}
Specifications
- Model
- novus-demo-1
- Tasks
- 5
- Metrics
- accuracy, f1, rougeL, bleu, exact_match
What is a .json file?
JSON (JavaScript Object Notation) is a lightweight, text-based data-interchange format representing objects, arrays, strings, numbers, booleans, and null. It is language-independent, human-readable, and the dominant format for web APIs and configuration. It requires a single well-formed root value.
How to use this file
Use an example JSON file to test parsers and serializers, schema validation, Unicode and number-precision handling, and API request or response processing.
How to use this file for testing
“Model Benchmark Results (JSON)” is a deterministic Novus Examples fixture for Model evaluation, JSON parsing, Data import. Benchmark results, confusion matrices, ROC curves, and classification reports in CSV and JSON — for testing eval dashboards, metric parsers, and leaderboard importers.
Documented properties for this file: JSON · 663 bytes. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
AI/ML fixtures are fully synthetic with documented schemas — no real people or data. Test data loaders, tokenizers, annotation converters, embedding/vector stores, or eval-metric parsers against the known structure and fixed seeds.
Code examples
import json
with open("benchmark-results.json") as f:
data = json.load(f)
print(type(data), len(data))Related files
- jsonClassification Report (JSON)A per-class classification report in the scikit-learn structure — precision, recall, F1, and support for each class plus accuracy and macro/weighted averages. A fixture for testing metric parsers and report renderers.

- jsonConfusion Matrix — 3 Classes (JSON)JSON twin of the 3-class confusion matrix.

- jsonConfusion Matrix — 3-class (JSON)The same 3-class confusion matrix as JSON — a labels array plus a nested counts matrix. The structured twin of the CSV, for testing evaluation tooling.

- jsonDetection Eval Metrics (JSON)Synthetic mAP evaluation summary for object-detection benchmark harness tests.

- jsonEval Metric — Accuracy MiniMinimal SAMPLE eval metric JSON (accuracy) for dashboard parsers.

- jsonEval Metric — Bleu MiniMinimal SAMPLE eval metric JSON (bleu) for dashboard parsers.

Generated by generation/ai_datasets.py. Free for any use, no attribution required — license.