Skip to content
Novus Examples
json751 B

Mini RAG Queries (JSON)

Retrieval-eval queries with relevant_docs pointers into the mini RAG Markdown corpus.

Preview — first 40 linesjson
{
  "queries": [
    {
      "id": "rq1",
      "query": "How long is the return window?",
      "relevant_docs": [
        "doc-a.md"
      ]
    },
    {
      "id": "rq2",
      "query": "When is Brightside Cafe open on weekends?",
      "relevant_docs": [
        "doc-b.md"
      ]
    },
    {
      "id": "rq3",
      "query": "Do I need to attribute Novus fixtures?",
      "relevant_docs": [
        "doc-c.md"
      ]
    },
    {
      "id": "rq4",
      "query": "What units should lab temperature use?",
      "relevant_docs": [
        "doc-d.md"
      ]
    },
    {
      "id": "rq5",
      "query": "Sample order id for returns?",
      "relevant_docs": [
        "doc-a.md"
      ]
    }
  ]
}

Specifications

Queries
5
Role
retrieval eval queries
Collection
mini-rag

What is a .json file?

JSON (JavaScript Object Notation) is a lightweight, text-based data-interchange format representing objects, arrays, strings, numbers, booleans, and null. It is language-independent, human-readable, and the dominant format for web APIs and configuration. It requires a single well-formed root value.

How to use this file

Use an example JSON file to test parsers and serializers, schema validation, Unicode and number-precision handling, and API request or response processing.

How to use this file for testing

“Mini RAG Queries (JSON)” is a deterministic Novus Examples fixture for ML training data, NLP datasets, JSON parsing. Labelled, synthetic datasets in the shapes ML pipelines expect — JSONL for text tasks, image annotations, embeddings, and sample weights — for testing data loaders, tokenizers, and training tooling.

Documented properties for this file: retrieval eval queries. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

AI/ML fixtures are fully synthetic with documented schemas — no real people or data. Test data loaders, tokenizers, annotation converters, embedding/vector stores, or eval-metric parsers against the known structure and fixed seeds.

These are labelled, training-shaped fixtures with a documented schema. Test your data loader, tokenizer, or format converter against it; every label and value is synthetic.

Code examples

import json

with open("queries.json") as f:
    data = json.load(f)
print(type(data), len(data))

Generated by generation/ai_wave_b.py. Free for any use, no attribution required — license.