How it works#
A SMILES string goes in. The engine reads the structure, applies the nomenclature rules, assembles a name and, for nearly every name, checks it with an OPSIN round trip before it leaves (How every name is checked). A few name classes that OPSIN cannot read leave on their construction alone; --provenance marks them.
Four stages do the work. The orchestrator, namer.py, runs them and holds the final OPSIN check.
Stage |
What it does |
|---|---|
Perception |
RDKit reads the structure; Orthonym finds rings, characteristic groups and stereocentres. The CIP descriptors come from the centres labeller; RDKit’s CIP labeller fills the few double bonds centres leaves unlabelled and takes over for a molecule centres cannot label. |
Rules |
Seniority, the parent hydride, locants, alphanumerical order and spelling, built as nomenclature classes from the IUPAC 2013 recommendations. Names the recommendations list one by one (retained and natural-product names), the names of metal tetrapyrrole complexes (ChEBI names, matched by exact InChIKey) and a last-resort table of trivial names are looked up by exact structure. |
Assembly |
Chains, rings, Hantzsch–Widman heterocycles, fused, bridged (von Baeyer) and spiro systems, and the characteristic-group families: acids, esters, amides, amines, nitriles and more. |
Validation |
The OPSIN round trip that decides whether a name leaves the engine, with named exceptions for classes OPSIN cannot read, and, for names from the general engine, an atom-coverage certificate. |
Compound classes are tried in a fixed order; a class that cannot build a name declines, and the next one is tried. There is no special case for a single molecule and no rewriting of an emitted name.
The source tree#
src/orthonym/
├── namer.py # orchestration and the final OPSIN check
├── cli.py # the command line
├── jars.py # finding, downloading and checking the OPSIN and centres jars
├── perception/ # structure perception: rings, characteristic groups, CIP stereo
├── routing/ # choosing the compound class
├── decomposition/ # naming large structures from their fragments
├── rules/ # IUPAC nomenclature rules
├── assembly/ # name assembly: locants, ordering, selection, the general engine
├── validation/ # OPSIN round-trip helpers and atom-coverage checks
├── metrics/ # provenance, tier labels and reason codes
└── data/ # naming tables
Every module has a page under Internals.
Built on#
RDKit, OPSIN and centres, on the IUPAC 2013 recommendations. The credits and references are in the README, under Built on.