Extractive QA Dataset — SQuAD v2 Format (JSON)
An extractive question-answering dataset in the SQuAD v2.0 JSON structure — titled articles with context paragraphs, questions, character-offset answers, and one deliberately unanswerable question. Synthetic content; a fixture for QA model training and SQuAD-format loaders.
{
"version": "v2.0",
"data": [
{
"title": "Natural Processes",
"paragraphs": [
{
"context": "The water cycle describes how water moves through the environment. Evaporation turns liquid water into vapour, which rises and cools to form clouds through condensation. Precipitation then returns the water to the surface as rain or snow.",
"qas": [
{
"id": "wc-001",
"question": "What turns liquid water into vapour?",
"answers": [
{
"text": "Evaporation",
"answer_start": 67
}
],
"is_impossible": false
},
{
"id": "wc-002",
"question": "What process forms clouds?",
"answers": [
{
"text": "condensation",
"answer_start": 156
}
],
"is_impossible": false
},
{
"id": "wc-003",
"question": "Which ocean current is mentioned?",
"answers": [],
"is_impossible": true
}
]
},
{
"context": "Photosynthesis allows plants to make food. Using sunlight, chlorophyll in the leaves converts carbon dioxide and water into glucose, releasing oxygen as a by-product.",
"qas": [
{
"id": "ps-001",
"question": "What do plants release as a by-product?",
"answers": [
{
"text": "oxygen",
"answer_start": 143
}Specifications
- Version
- v2.0
- Questions
- 5
- Has Impossible
- true
- Task
- extractive question answering
What is a .json file?
JSON (JavaScript Object Notation) is a lightweight, text-based data-interchange format representing objects, arrays, strings, numbers, booleans, and null. It is language-independent, human-readable, and the dominant format for web APIs and configuration. It requires a single well-formed root value.
How to use this file
Use an example JSON file to test parsers and serializers, schema validation, Unicode and number-precision handling, and API request or response processing.
How to use this file for testing
“Extractive QA Dataset — SQuAD v2 Format (JSON)” is a deterministic Novus Examples fixture for ML training data, NLP datasets, JSON parsing. Labelled, synthetic datasets in the shapes ML pipelines expect — JSONL for text tasks, image annotations, embeddings, and sample weights — for testing data loaders, tokenizers, and training tooling.
Documented properties for this file: task: extractive question answering. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
AI/ML fixtures are fully synthetic with documented schemas — no real people or data. Test data loaders, tokenizers, annotation converters, embedding/vector stores, or eval-metric parsers against the known structure and fixed seeds.
These are labelled, training-shaped fixtures with a documented schema. Test your data loader, tokenizer, or format converter against it; every label and value is synthetic.
Code examples
import json
with open("qa-squad.json") as f:
data = json.load(f)
print(type(data), len(data))Related files
- jsonlChat Fine-tuning Dataset — Anthropic Format (JSONL)The same synthetic conversations in the Anthropic Messages JSONL shape — a top-level system prompt plus a messages array of user and assistant turns. The format twin of the OpenAI file, for testing chat-format conversion.

- jsonlChat Fine-tuning Dataset — OpenAI Format (JSONL)A chat fine-tuning dataset in the OpenAI JSONL format — one conversation per line as a messages array with system, user, and assistant turns. Synthetic Q&A content. Paired with an Anthropic-format twin for testing format converters.

- jsonlClassification Dataset — Hierarchical (JSONL)Hierarchical category labels for taxonomy-aware classifiers.

- jsonlClassification Dataset — Imbalanced (JSONL)Imbalanced label distribution — 2 rare vs 8 common rows.

- jsonlClassification Dataset — Intent (JSONL)Intent classification JSONL for dialog systems.

- jsonlClassification Dataset — Language Id (JSONL)Language identification JSONL with EN/ES pairs.

Generated by generation/ai_datasets.py. Free for any use, no attribution required — license.