Skip to content
Novus Examples
jsonl2.1 KB

Instruction-tuning Dataset — Alpaca Format (JSONL)

An instruction-tuning dataset in the Alpaca JSONL format — 20 instruction / input / output triples covering small transformations, extraction, and factual answers. Fully synthetic; a fixture for supervised fine-tuning pipelines.

Preview — first 21 linesjsonl
{"instruction": "Summarize the text in one sentence.", "input": "The library closes at 8pm on weekdays and 5pm on weekends.", "output": "The library closes at 8pm on weekdays and 5pm on weekends."}
{"instruction": "Translate to French.", "input": "Good morning, how are you?", "output": "Bonjour, comment allez-vous ?"}
{"instruction": "Classify the sentiment as positive or negative.", "input": "I absolutely loved the concert.", "output": "positive"}
{"instruction": "Extract the email address.", "input": "Reach me at dev@example.com after noon.", "output": "dev@example.com"}
{"instruction": "Convert to uppercase.", "input": "hello world", "output": "HELLO WORLD"}
{"instruction": "List three primary colours.", "input": "", "output": "Red, blue, and yellow."}
{"instruction": "What is 15 percent of 200?", "input": "", "output": "30"}
{"instruction": "Correct the grammar.", "input": "She don't like apples.", "output": "She doesn't like apples."}
{"instruction": "Give the plural form.", "input": "mouse", "output": "mice"}
{"instruction": "Sort the numbers ascending.", "input": "5, 2, 9, 1", "output": "1, 2, 5, 9"}
{"instruction": "Name the largest planet in the solar system.", "input": "", "output": "Jupiter"}
{"instruction": "Rewrite the sentence in the past tense.", "input": "I walk to school.", "output": "I walked to school."}
{"instruction": "Provide the chemical symbol for gold.", "input": "", "output": "Au"}
{"instruction": "Count the words.", "input": "the quick brown fox", "output": "4"}
{"instruction": "Turn this into a question.", "input": "The store is open.", "output": "Is the store open?"}
{"instruction": "Give an antonym for 'increase'.", "input": "", "output": "decrease"}
{"instruction": "Round to the nearest whole number.", "input": "3.7", "output": "4"}
{"instruction": "Identify the language.", "input": "Hola, ¿cómo estás?", "output": "Spanish"}
{"instruction": "Extract the year.", "input": "The treaty was signed in 1848.", "output": "1848"}
{"instruction": "Abbreviate the phrase.", "input": "as soon as possible", "output": "ASAP"}

Specifications

Records
20
Format
Alpaca
Schema
instruction, input, output
Task
instruction tuning

What is a .jsonl file?

JSONL (JSON Lines) is a text format where each line is a complete, independent JSON value, allowing records to be streamed and appended without parsing the whole file. It is not itself a JSON array and each line must stand alone. It is common in logging, machine learning datasets, and data pipelines.

How to use this file

Use an example JSONL to test line-by-line streaming parsers, append-and-resume ingestion, and batch pipelines that process one record per line.

How to use this file for testing

“Instruction-tuning Dataset — Alpaca Format (JSONL)” is a deterministic Novus Examples fixture for ML training data, NLP datasets, JSON parsing. Labelled, synthetic datasets in the shapes ML pipelines expect — JSONL for text tasks, image annotations, embeddings, and sample weights — for testing data loaders, tokenizers, and training tooling.

Documented properties for this file: 20 records · schema: instruction, input, output · task: instruction tuning. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

AI/ML fixtures are fully synthetic with documented schemas — no real people or data. Test data loaders, tokenizers, annotation converters, embedding/vector stores, or eval-metric parsers against the known structure and fixed seeds.

These are labelled, training-shaped fixtures with a documented schema. Test your data loader, tokenizer, or format converter against it; every label and value is synthetic.

Code examples

import json

with open("instructions.jsonl") as f:
    rows = [json.loads(line) for line in f]
print(len(rows), rows[0])

Generated by generation/ai_datasets.py. Free for any use, no attribution required — license.