Skip to content
Novus Examples
csv512 B

CLDR Plural-Category Matrix — six languages

One row per language, one column per CLDR plural category, filled with the smallest integers that select it — the sheet to hand a translator or a PM when explaining why 'one' and 'other' is not enough. Japanese needs one form and Welsh needs six.

Preview — first 8 linescsv
language,locale,cardinalCategories,zero,one,two,few,many,other,gettextNplurals
English,en,2,(unused),1,(unused),(unused),(unused),"0, 2, 3, 100",2
Japanese,ja,1,(unused),(unused),(unused),(unused),(unused),"0, 1, 2, 5",1
Polish,pl,4,(unused),1,(unused),"2, 3, 4, 22","0, 5, 9, 10","0.5, 1.0, 1.5, 2.0",3
Russian,ru,4,(unused),"1, 21, 31, 101",(unused),"2, 3, 4, 22","0, 5, 6, 9","0.5, 1.0, 1.5, 10.0",3
Arabic,ar,6,0,1,2,"3, 4, 9, 10","11, 26, 99, 111","100, 101, 102, 200",6
Welsh,cy,6,0,1,2,3,6,"4, 5, 7, 8",6

Specifications

Rows
6
Columns
10
Locales
en, ja, pl, ru, ar, cy
Category Range
1 to 6
Delimiter
,
Header
true
Seed
20260807
Wave
p7
Line Endings
LF

Testing contract

Reference control
Scenario
Join this table against your message catalog to find keys that ship fewer plural forms than the target language requires.
Expected result
Every '(unused)' cell marks a category the language never selects; every filled cell is a count your catalog must have a form for.

What is a .csv file?

CSV (Comma-Separated Values) is a plain-text tabular format where rows are lines and fields are separated by commas, with quoting rules for values that contain delimiters, quotes, or newlines. It has no formal type system and depends on encoding and dialect conventions. It is the most portable format for tabular data exchange.

How to use this file

Use an example CSV to test parsers against quoting and embedded-delimiter edge cases, header handling, encoding detection, and import pipelines into databases or spreadsheets.

How to use this file for testing

“CLDR Plural-Category Matrix — six languages” is a deterministic Novus Examples fixture for Internationalization, Localization catalogs, CSV parsing. Parallel text and accented, multi-script content — for testing translation pipelines, Unicode handling, and localization tooling.

Documented properties for this file: seed 20260807 · 6 rows · 10 columns · LF. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

Translation-catalog fixtures carry the same message set across formats, each with its native placeholder syntax. Test your i18n loader, catalog converter, or translation-memory tool, and use the RTL and CJK variants to check bidirectional text and Unicode handling.

Load the catalog with your i18n framework and verify placeholder interpolation and plural handling; the RTL and CJK variants exercise bidirectional text and font fallback.

Code examples

import pandas as pd

df = pd.read_csv("plural-category-matrix.csv")
print(df.head())
print(df.dtypes)

Generated by generation/localization_p7.py. Free for any use, no attribution required — license.