Skip to content
Novus Examples
smi12 B

Ethanol — SMILES String (.smi)

Ethanol as a single tab-delimited SMILES record, the most compressed member of this family: three heavy atoms with every hydrogen implicit and no geometry whatsoever. Round-tripping molfile to SMILES and back is the classic lossy conversion, and this pair is the reference for it.

Preview — first 2 linessmi
CCO	ethanol

Specifications

Format
SMILES
Records
1
Smiles
CCO
Delimiter
tab
Heavy Atoms
3
Hydrogens
implicit
Coordinates
none

Testing contract

Expected to pass
Scenario
Parse the SMILES string, add explicit hydrogens, and compare the resulting formula with the paired molfile.
Expected result
The molecular formula comes out as C2H6O with nine atoms after hydrogens are added, matching the molfile's atom count while carrying none of its coordinates.

What is a .smi file?

A .smi file holds SMILES strings — a line notation that encodes a molecular graph as text. Atoms are written as element symbols, aromatic atoms in lower case, bonds as `-`, `=`, `#`, branches in parentheses, and rings as matching digit labels, with stereochemistry expressed by `/`, `\\`, and `@` markers. Files typically carry one SMILES per line with an optional whitespace-separated identifier, and a canonical SMILES is a unique string for a given structure.

How to use this file

Use an example .smi file to test SMILES parsers, canonicalisers, and structure-search tooling, verifying ring-closure and aromaticity handling, stereochemistry round-tripping, and that an invalid string is rejected rather than partially parsed.

How to use this file for testing

“Ethanol — SMILES String (.smi)” is a deterministic Novus Examples fixture for Scientific data, Editor testing, Conversion testing. Citation catalogs (BibTeX, RIS), chemistry structures (MDL Molfile, PDB), and gridded binary data (NetCDF, FITS) — for testing reference managers, molecule viewers, and scientific-data loaders.

Documented properties for this file: 1 records · SMILES. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.

Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such — expect parsers to fail loudly rather than silently accept them.

Scientific fixtures are small, valid, and fully synthetic — no real organism, patient, sample, or observation. Point your parser or loader at the file and check it reads the documented records, variables, or headers; binary formats ship a readable twin or metadata listing for comparison.

Generated by generation/scientific.py. Free for any use, no attribution required — license.