Sentiment Classification Dataset (JSONL)
A labelled sentiment-classification dataset in JSON Lines — 24 short product-review-style sentences balanced across positive, negative, and neutral. Fully synthetic; a fixture for testing text-classification loaders, tokenizers, and JSONL parsers.
{"text": "The battery lasts all day and the screen is gorgeous.", "label": "positive"}
{"text": "Arrived two weeks late and the box was crushed.", "label": "negative"}
{"text": "It works as described. Nothing surprising either way.", "label": "neutral"}
{"text": "Best purchase I've made this year — highly recommend.", "label": "positive"}
{"text": "Stopped charging after a month. Very disappointed.", "label": "negative"}
{"text": "Setup took a while but support was helpful.", "label": "neutral"}
{"text": "Incredibly comfortable and the build quality feels premium.", "label": "positive"}
{"text": "The app crashes every time I open the settings page.", "label": "negative"}
{"text": "Does the job. Fairly average for the price.", "label": "neutral"}
{"text": "Fast shipping and exactly what I ordered.", "label": "positive"}
{"text": "Instructions were confusing and parts were missing.", "label": "negative"}
{"text": "Fine for casual use, not for anything demanding.", "label": "neutral"}
{"text": "The sound is crisp and the bass is well balanced.", "label": "positive"}
{"text": "Overpriced for what you actually get.", "label": "negative"}
{"text": "Looks nice on the desk; performance is unremarkable.", "label": "neutral"}
{"text": "Customer service replaced it without any hassle.", "label": "positive"}
{"text": "It overheats within minutes of heavy use.", "label": "negative"}
{"text": "Standard packaging, standard product, no complaints.", "label": "neutral"}
{"text": "Lightweight, sturdy, and surprisingly affordable.", "label": "positive"}
{"text": "The colours look nothing like the photos online.", "label": "negative"}
{"text": "Reasonable quality but the manual could be clearer.", "label": "neutral"}
{"text": "Exceeded my expectations in every way.", "label": "positive"}
{"text": "Broke on the second day. Would not buy again.", "label": "negative"}
{"text": "Perfectly adequate for everyday tasks.", "label": "neutral"}
Specifications
- Records
- 24
- Labels
- positive, negative, neutral
- Schema
- text, label
- Task
- text classification
What is a .jsonl file?
JSONL (JSON Lines) is a text format where each line is a complete, independent JSON value, allowing records to be streamed and appended without parsing the whole file. It is not itself a JSON array and each line must stand alone. It is common in logging, machine learning datasets, and data pipelines.
How to use this file
Use an example JSONL to test line-by-line streaming parsers, append-and-resume ingestion, and batch pipelines that process one record per line.
How to use this file for testing
“Sentiment Classification Dataset (JSONL)” is a deterministic Novus Examples fixture for ML training data, NLP datasets, JSON parsing, Data import. Labelled, synthetic datasets in the shapes ML pipelines expect — JSONL for text tasks, image annotations, embeddings, and sample weights — for testing data loaders, tokenizers, and training tooling.
Documented properties for this file: 24 records · schema: text, label · task: text classification. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
AI/ML fixtures are fully synthetic with documented schemas — no real people or data. Test data loaders, tokenizers, annotation converters, embedding/vector stores, or eval-metric parsers against the known structure and fixed seeds.
These are labelled, training-shaped fixtures with a documented schema. Test your data loader, tokenizer, or format converter against it; every label and value is synthetic.
Code examples
import json
with open("sentiment.jsonl") as f:
rows = [json.loads(line) for line in f]
print(len(rows), rows[0])Related files
- jsonlChat Fine-tuning Dataset — Anthropic Format (JSONL)The same synthetic conversations in the Anthropic Messages JSONL shape — a top-level system prompt plus a messages array of user and assistant turns. The format twin of the OpenAI file, for testing chat-format conversion.

- jsonlChat Fine-tuning Dataset — OpenAI Format (JSONL)A chat fine-tuning dataset in the OpenAI JSONL format — one conversation per line as a messages array with system, user, and assistant turns. Synthetic Q&A content. Paired with an Anthropic-format twin for testing format converters.

- jsonlClassification Dataset — Hierarchical (JSONL)Hierarchical category labels for taxonomy-aware classifiers.

- jsonlClassification Dataset — Imbalanced (JSONL)Imbalanced label distribution — 2 rare vs 8 common rows.

- jsonlClassification Dataset — Intent (JSONL)Intent classification JSONL for dialog systems.

- jsonlClassification Dataset — Language Id (JSONL)Language identification JSONL with EN/ES pairs.

Generated by generation/ai_datasets.py. Free for any use, no attribution required — license.