Ethanol — SDfile With Data Fields (.sdf)
An SDfile wrapping the identical ethanol molfile plus five tagged data fields and the mandatory $$$$ terminator. The property block syntax — a header line, a value and a blank line — is where SDfile parsers usually diverge from molfile parsers.
ethanol
NovusEx 3D synthetic idealised geometry
Idealised geometry. Property block follows the connection table.
9 8 0 0 0 0 0 0 0 0999 V2000
-1.2456 0.2436 0.0000 C 0 0 0 0 0 0 0 0 0 0 0 0
0.0000 -0.5876 0.0000 C 0 0 0 0 0 0 0 0 0 0 0 0
1.1636 0.2296 0.0000 O 0 0 0 0 0 0 0 0 0 0 0 0
-2.1345 -0.3861 0.0000 H 0 0 0 0 0 0 0 0 0 0 0 0
-1.2734 0.8797 0.8837 H 0 0 0 0 0 0 0 0 0 0 0 0
-1.2734 0.8797 -0.8837 H 0 0 0 0 0 0 0 0 0 0 0 0
0.0384 -1.2270 0.8850 H 0 0 0 0 0 0 0 0 0 0 0 0
0.0384 -1.2270 -0.8850 H 0 0 0 0 0 0 0 0 0 0 0 0
1.9524 -0.3175 0.0000 H 0 0 0 0 0 0 0 0 0 0 0 0
1 2 1 0
2 3 1 0
1 4 1 0
1 5 1 0
1 6 1 0
2 7 1 0
2 8 1 0
3 9 1 0
M END
> <NAME>
ethanol
> <FORMULA>
C2H6O
> <MOLECULAR_WEIGHT>
46.069
> <SMILES>
CCO
> <SOURCE>
Novus Examples synthetic fixture (idealised geometry)
$$$$
Specifications
- Format
- MDL SDfile
- Records
- 1
- Data Fields
- 5
- Record Terminator
- $$$$
- Field Syntax
- > <NAME> then value then a blank line
- Embedded Molfile
- identical to the paired .mol
Testing contract
Expected to pass- Scenario
- Parse the record, extract the embedded connection table and read all five tagged data fields.
- Expected result
- The connection table is byte-identical to the paired molfile, the five fields parse with FORMULA equal to C2H6O, and the record terminates on $$$$ rather than at end of file.
What is a .sdf file?
SDF (Structure-Data File) is MDL's multi-record chemistry format. Each record is a molfile — a counts line, an atom block, and a bond block — followed by tagged data fields written as `> <NAME>` lines with their values, and terminated by a `$$$$` delimiter. It is how compound libraries with per-molecule properties are shipped between cheminformatics tools.
How to use this file
Use an example .sdf file to test cheminformatics readers and property extractors, verifying record splitting on the `$$$$` delimiter, that data fields are associated with the right structure, and that a malformed record does not consume the rest of the file.
How to use this file for testing
“Ethanol — SDfile With Data Fields (.sdf)” is a deterministic Novus Examples fixture for Scientific data, Editor testing, Conversion testing. Citation catalogs (BibTeX, RIS), chemistry structures (MDL Molfile, PDB), and gridded binary data (NetCDF, FITS) — for testing reference managers, molecule viewers, and scientific-data loaders.
Documented properties for this file: 1 records · MDL SDfile. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.
Scientific fixtures are small, valid, and fully synthetic — no real organism, patient, sample, or observation. Point your parser or loader at the file and check it reads the documented records, variables, or headers; binary formats ship a readable twin or metadata listing for comparison.
Related files
- csvChemistry Fixture Molecule Index (.csv)One row per molecule used across the chemistry fixtures, with formula, molecular weight, SMILES and the atom and bond counts each file should yield. It is the oracle a cheminformatics toolkit can be scored against without needing a second toolkit to generate the answers.

- csvDICOM-Shaped Dataset — Element Table Reference (.csv)Every element of the synthetic phantom dataset as tag, value representation, meaning and encoded length, in the ascending tag order the standard mandates. Diff a parser's element list against it to check both the ordering rule and the two different length encodings.

- fastaFASTA Contigs — One Line Per Record (.fasta)The same four contigs written one sequence per line rather than wrapped, which is how many pipelines emit FASTA and how many line-oriented parsers assume it always looks. Comparing against the wrapped twin proves a parser joins continuation lines instead of taking the first one.

- csvFASTQ Quality Score Reference — Both Encodings Decoded (.csv)Every twelfth base of the six reads with its Phred score, the ASCII character it takes under both offsets, and the error probability that score implies. It converts a quality-encoding argument into a lookup you can diff.

- fastqFASTQ Reads — Phred+64 Legacy Encoding (.fastq)The identical reads and identical quality scores written with the legacy Phred+64 offset used by older Illumina pipelines. Decode it with the modern offset and every base looks 31 points better than it is, which is a silent quality inflation rather than a parse failure.

- csvFITS WCS Header Cards — CSV Reference (.csv)Every header card of the tangent-plane WCS file transcribed to keyword, value and comment columns. Diff a header parser's output against it to prove the parser split each 80-column card at the right places instead of guessing on whitespace.

Generated by generation/scientific.py. Free for any use, no attribution required — license.