Skip to content
Novus Examples
md145 B

Mini RAG Corpus — doc-c.md

A short Markdown document in the mini RAG corpus (doc-c.md) — retrieve against the paired queries JSON.

Preview — first 4 linesmd
# Novus Examples License

All fixtures are free for any use with no attribution required. Corrupt samples are labelled intentionally corrupt.

Specifications

Role
corpus document
Collection
mini-rag
Format
Markdown

What is a .md file?

Markdown (MD) is a lightweight plain-text markup language that uses simple punctuation conventions to denote headings, lists, links, emphasis, and code. It is designed to be readable as-is and to convert cleanly to HTML. It is widely used for documentation, READMEs, and content authoring.

How to use this file

Use an example Markdown file to test parsers and renderers, verify GitHub-Flavored Markdown extensions like tables and fenced code, and exercise HTML-conversion pipelines.

How to use this file for testing

“Mini RAG Corpus — doc-c.md” is a deterministic Novus Examples fixture for ML training data, NLP datasets, Editor testing. Labelled, synthetic datasets in the shapes ML pipelines expect — JSONL for text tasks, image annotations, embeddings, and sample weights — for testing data loaders, tokenizers, and training tooling.

Documented properties for this file: corpus document · Markdown. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

AI/ML fixtures are fully synthetic with documented schemas — no real people or data. Test data loaders, tokenizers, annotation converters, embedding/vector stores, or eval-metric parsers against the known structure and fixed seeds.

These are labelled, training-shaped fixtures with a documented schema. Test your data loader, tokenizer, or format converter against it; every label and value is synthetic.

Code examples

import markdown  # pip install markdown

html = markdown.markdown(open("doc-c.md").read())
print(html[:200])

Generated by generation/ai_wave_b.py. Free for any use, no attribution required — license.