Chat Fine-tuning Dataset — OpenAI Format (JSONL)
A chat fine-tuning dataset in the OpenAI JSONL format — one conversation per line as a messages array with system, user, and assistant turns. Synthetic Q&A content. Paired with an Anthropic-format twin for testing format converters.
{"messages": [{"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What's the capital of Australia?"}, {"role": "assistant", "content": "The capital of Australia is Canberra, not Sydney as many people assume."}]}
{"messages": [{"role": "system", "content": "You are a concise coding assistant."}, {"role": "user", "content": "How do I reverse a list in Python?"}, {"role": "assistant", "content": "Use slicing: `reversed_list = my_list[::-1]`, or call `my_list.reverse()` to reverse it in place."}]}
{"messages": [{"role": "system", "content": "You are a friendly tutor."}, {"role": "user", "content": "Explain photosynthesis in one sentence."}, {"role": "assistant", "content": "Photosynthesis is how plants convert sunlight, water, and carbon dioxide into glucose and oxygen."}]}
{"messages": [{"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Convert 25 degrees Celsius to Fahrenheit."}, {"role": "assistant", "content": "25°C equals 77°F. The formula is (C × 9/5) + 32."}]}
{"messages": [{"role": "system", "content": "You are a travel guide."}, {"role": "user", "content": "What should I pack for a rainy city break?"}, {"role": "assistant", "content": "A compact umbrella, a waterproof jacket, quick-dry shoes, and a dry bag for electronics."}]}
{"messages": [{"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is a prime number?"}, {"role": "assistant", "content": "A prime number is a whole number greater than 1 whose only divisors are 1 and itself, like 2, 3, 5, and 7."}]}
{"messages": [{"role": "system", "content": "You are a cooking assistant."}, {"role": "user", "content": "How long should I boil an egg for it to be hard-boiled?"}, {"role": "assistant", "content": "About 9 to 12 minutes in boiling water, then transfer it to cold water to stop cooking."}]}
{"messages": [{"role": "system", "content": "You are a concise assistant."}, {"role": "user", "content": "Give me a synonym for 'happy'."}, {"role": "assistant", "content": "'Cheerful' works well; 'content' and 'delighted' are good alternatives depending on the tone."}]}
{"messages": [{"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What does HTTP stand for?"}, {"role": "assistant", "content": "HTTP stands for HyperText Transfer Protocol, the foundation of data exchange on the web."}]}
{"messages": [{"role": "system", "content": "You are a fitness coach."}, {"role": "user", "content": "How many steps a day is often recommended?"}, {"role": "assistant", "content": "A common target is around 8,000 to 10,000 steps a day, though any increase over your baseline helps."}]}
Specifications
- Conversations
- 10
- Format
- OpenAI messages
- Roles
- system, user, assistant
- Task
- chat fine-tuning
What is a .jsonl file?
JSONL (JSON Lines) is a text format where each line is a complete, independent JSON value, allowing records to be streamed and appended without parsing the whole file. It is not itself a JSON array and each line must stand alone. It is common in logging, machine learning datasets, and data pipelines.
How to use this file
Use an example JSONL to test line-by-line streaming parsers, append-and-resume ingestion, and batch pipelines that process one record per line.
How to use this file for testing
“Chat Fine-tuning Dataset — OpenAI Format (JSONL)” is a deterministic Novus Examples fixture for ML training data, NLP datasets, JSON parsing. Labelled, synthetic datasets in the shapes ML pipelines expect — JSONL for text tasks, image annotations, embeddings, and sample weights — for testing data loaders, tokenizers, and training tooling.
Documented properties for this file: task: chat fine-tuning · OpenAI messages. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
AI/ML fixtures are fully synthetic with documented schemas — no real people or data. Test data loaders, tokenizers, annotation converters, embedding/vector stores, or eval-metric parsers against the known structure and fixed seeds.
These are labelled, training-shaped fixtures with a documented schema. Test your data loader, tokenizer, or format converter against it; every label and value is synthetic.
Code examples
import json
with open("chat-openai.jsonl") as f:
rows = [json.loads(line) for line in f]
print(len(rows), rows[0])Related files
- jsonlClassification Dataset — Hierarchical (JSONL)Hierarchical category labels for taxonomy-aware classifiers.

- jsonlClassification Dataset — Imbalanced (JSONL)Imbalanced label distribution — 2 rare vs 8 common rows.

- jsonlClassification Dataset — Intent (JSONL)Intent classification JSONL for dialog systems.

- jsonlClassification Dataset — Language Id (JSONL)Language identification JSONL with EN/ES pairs.

- jsonlClassification Dataset — Multilabel (JSONL)Multi-label classification JSONL with two labels per row.

- jsonlClassification Dataset — Topic (JSONL)Topic classification rows for news routing tests.

Generated by generation/ai_datasets.py. Free for any use, no attribution required — license.