Changelog#
All notable changes to Orthonym are documented here.
The format follows Keep a Changelog and this project adheres to Semantic Versioning.
1.0.3 (2026-10-01)#
Bug Fixes#
bridged fused PINs on naphthalene and anthracene parents, larger ring systems at the best-effort tier, more PIN classes and exact spellings (7bdb0d2)
1.0.2 (2026-09-30)#
Bug Fixes#
engine fixes and optimizations, tier updates (6cab387)
1.0.1 (2026-09-29)#
The naming engine is the same as in 1.0.0: the Python code differs only in comments and docstrings. The PyPI package now lists all three authors and Development Status 5 - Production/Stable, and the documentation cites the Zenodo archive.
Documentation#
the Zenodo DOI in the README, the docs and CITATION.cff (9415508)
1.0.0 (2026-09-29)#
Orthonym 1.0.0 is the first public release of a deterministic, rule-based generator that turns a SMILES string into an IUPAC name, aiming at the Preferred IUPAC Name (PIN) of the IUPAC 2013 recommendations. Each name is parsed back by OPSIN and must match the input structure by full InChIKey before it is shown; otherwise Orthonym declines and gives the reason, and the few names OPSIN cannot read in full are labelled as such. The release gathers 2,799 public commits made from late January to 29 September 2026, which took the engine from chains, monocycles and the main functional groups to fused, bridged, spiro and phane ring systems, stereodescriptors, charged species, isotopes, natural products, peptides, glycans, lipids and metal compounds. On the corpora of the accompanying paper, run with the paper’s recipe on the release branch, it gave no wrong names.
Highlights#
Rules only, with no neural network and no sampling: the same SMILES always gives the same name.
Names are read back by OPSIN and compared with the input in constitution, charge and stereo; a name that fails is withdrawn, and when none is left Orthonym declines with a reason.
A PIN tier for names built and certified on the strict preferred-name path, and opt-in wider tiers up to best effort (
--emit-tier) whose names must pass a full-InChIKey round trip (metal-complex list names excepted).--provenanceprints each result as JSON with its tier, its source and averifiedfield that says how the name was checked.0 wrong names on the paper’s corpora (measured on the release branch with the paper’s recipe); round-trip exact: QM9 133,860 of 133,885, ChEBI 107,946 of 111,843, PubChem 500k 499,666 of 500,000, ZINC22 500k 488,428 of 500,000.
Ring systems: monocycles, Hantzsch-Widman heterocycles, fused systems with IUPAC numbering, von Baeyer polycycles, spiro compounds, phanes and ring assemblies.
Stereo: CIP R/S from the centres engine, E/Z, ring cis/trans and pseudoasymmetric r/s, each checked against the input’s CIP labels.
Ions, salts, zwitterions, radicals, isotope-labelled compounds, natural products, peptides, oligosaccharides, lipids, nucleotides and organometallics.
OPSIN runs in the same process through JPype; the OPSIN and centres jars are downloaded at install and checked against pinned SHA-256 checksums.
Documentation at https://steinbeck-lab.github.io/Orthonym/ and a web app at https://orthonym.decimer.ai (source: https://github.com/Steinbeck-Lab/Orthonym-Web).
Nomenclature coverage#
Chains and substituents
Acyclic chains with IUPAC numerical terms, lowest locants and double- and triple-bond locants (956329481), including unbranched alkanes of 80 carbons and more (68f50d128).
One substituent-naming path shared by the ester, lactone, lactam, amide, acid-halide, anhydride and benzene namers, so substituents are not silently dropped (08b6e0d62); alkenyl and alkynyl prefixes (d63e482b3).
Nested enclosing marks for compound substituents, including N-substituents on amides (P-16.5) (eb5d40c69); alphabetical tie-breaking of chain orientation (ec90679f5).
Multiplicative names (bis/tris/tetrakis) that keep every identical arm, with bridges such as sulfanediyl, peroxy and disulfanediyl (bcea5da52, 5b2fc9c87).
Parent selection and functional groups
The parent chain or ring is chosen by the P-44 criteria in order: principal groups, ring or chain, length, multiple bonds, lowest locants (863ff10db, 6cf607ac7).
Suffix seniority follows the functional-group order of P-41, and lower-ranked groups are cited as prefixes (4b6158840, d396a18c9).
Nitriles, amides, esters, acid halides, anhydrides, hydroperoxides, carbamic and thiocarboxylic acids, and functional class names for oximes, N-oxides, isocyanates and carbamates (b73e049cf).
Seleno and telluro acids (24fb4c019), sulfinamides (62bd3f515), amidines, imidates, carbamoylamino and ylidene prefixes, and principal chains of As, Sb, Se and Te oxoacids.
Rings: monocycles, Hantzsch-Widman and fused systems
Cycloalkanes, cycloalkenes and Hantzsch-Widman names for 3- to 10-membered heterocycles, such as oxolane and thiazolidine (785a1e0c3); unsaturated 7- and 8-membered heterocycles take Hantzsch-Widman names (c6599a59b).
Retained fused heterocycles with IUPAC peripheral numbering (87146be5e); fused systems not in the dictionary are named from their components, the base component chosen by P-25.3.1.3 (a821b6b53).
A fusion-numbering engine for cata-fused arenes and mixed 5/6-membered systems, which declines ring systems it does not cover (6a82e08bf).
Substituted two-component ortho-fused heterocycles and naphtho-fused systems (bb07d4ba6, a2a12971b).
Indicated hydrogen, added indicated hydrogen and hydro prefixes, with cyclic ketones and quinones named as PIN diones (7472e8dca); chromones and coumarins on the 1-benzopyran parent (757586d58).
Rings: von Baeyer, spiro, phanes and ring assemblies
Von Baeyer names with secondary and zero-atom bridges and heteroatom replacement (5fde4e994); main ring and main bridge chosen for PIN descriptors, preferring the largest main bridge (ce8a95d7b).
Spiro names, including dispiro and heterospiro systems, spiro compounds of von Baeyer and fluorene components (e991c19c7), and dispiro and trispiro systems of fused components (8e66cc181).
Preferred names for monocyclic all-benzene homophanes (a3c878a2a); ring assemblies with primes (1,1’:4’,1’’-terphenyl), enclosing marks and multipliers up to deci.
Stereochemistry
R/S and E/Z with locants (6bb4693a9), with the centres engine as the default source of CIP labels (0537e1517).
Descriptors on substituents, oximes, esters, lactones, lactams, fused heterocycles and polycycles, with CIP labels recomputed on each fragment (646fe501a, 0179bca74).
Chain and ring stereocentres expressed in full; a name that leaves out or invents a stereocentre is rejected (c242eb2c7); cis/trans on rings with two stereocentres (9288fd0e3).
Steroid alpha/beta descriptors (e4056e902).
Charged species, salts, zwitterions and radicals
Each input is classified as neutral, ion, zwitterion, salt or radical (867cc6026), and salts are named as cation plus anion (512aed128).
Semipolar oxides (oxido/-ium), azide, diazo, nitrooxy and isocyano groups (2803b3a97); nitro and N-oxide groups are named as neutral groups (37461389e).
Stereo ionic salts, protonated diamines, hemisalts and mixed multi-anion salts; a salt with a part that cannot be named is refused (2c622fff5).
Carbon, chalcogen, ring-N, N-oxyl, acyl, nitrene and multicentre radicals, each shown only after a round trip that compares the radical sites (34ef0d3af).
Isotopes
Isotopic descriptors ordered by symbol and mass, with counts as subscripts, as in (1,1-2H2) (50331d9a9); labelled retained names fall back to a systematic parent, and labels go on the right word of a salt name (f5eaa886a).
Natural products and retained names
Steroid and alkaloid scaffolds named from the scaffold, with IUPAC atom numbering for nine steroid skeletons (55ed43bb3); nor-, homo- and seco- modifications (c1bc100a2).
Terpene stereoparents (IUPAC Table 10.1c), a complete D/L set of aldonic acids and xylitol (7b1175c72); ring stereo on saturated steroid scaffolds (815c950f6).
Retained names that are not PINs give way to systematic ones (acetone becomes propan-2-one, picric acid 2,4,6-trinitrophenol); retained names from OPSIN data are round-trip checked before use (163007b89, 5a696fd61).
Peptides, glycans, lipids and nucleotides
Peptides as acylamino chains, including capped termini, ester C-termini and non-standard backbones (6515385de, d5371532a).
An amino acid with a substituent on its nitrogen takes its systematic name, e.g. (2S)-1-acetylpyrrolidine-2-carboxylic acid (68f50d128).
About 50 retained sugar names and glycosides (769783748); disaccharides and oligosaccharides, including non-reducing ones such as raffinose (419c678b8).
Triglycerides and polyol polyesters (6a85190b4); phosphatidylcholine, phosphatidylethanolamine, phosphatidylserine, phosphatidic acid, ceramides and sphingolipids (9db641d5f).
Nucleosides and nucleotides with modified sugars or substituted bases (d64ccb421); acyl-CoA molecules on their 9H-purine parent (9f236276e).
Organometallics and inorganic parents
Metallocenes such as ferrocene, mononuclear metal carbonyls and half-sandwich complexes (9e0fe162a); sigma-bonded main-group organometallics (95ef27193).
Additive names for sigma-coordinated and metallacycle compounds; compounds with two metals are refused rather than guessed (a147f99a1).
An exact-InChIKey table of retained coordination names drawn from ChEBI, including corrinoid precursors (c75c88549), and hydrate word forms (mono, di, hemi, sesqui).
Parent hydrides of Group 15, the chalcogens, the halogens, boron and Group 14, polyazanes and lambda-convention hydrides (376445ebf).
Large molecules and mixtures
Large molecules split at ester, amide, ether, glycosidic, thioester, phosphodiester and sulfonamide bonds, including several bonds of mixed types, with each piece named and the name rebuilt (26c75704c, c2f90e2d0).
Neutral multi-component SMILES, such as cocrystals, named one component at a time in a fixed order (f43b8e3e4).
Checks and correctness#
Every public entry point checks its own names by an OPSIN round trip that compares constitution, charge and stereo (0be43e0e5); a name that OPSIN reads as a different molecule is not shown (5f43cd2bf).
The default tier also names a few classes OPSIN cannot read in full: names from exact-match lists (metal-complex and natural-product parent names), a few name forms OPSIN’s grammar lacks or misreads, and names whose stereodescriptors OPSIN cannot parse (constitution confirmed by OPSIN, each descriptor checked against its CIP label).
--provenancemarks each; the wider tiers ship none of them except the metal-complex list names.Tiers reported by
--provenance:pin_verified,pin_unverified,systematic_verified,best_effortandabstain;verifiedisopsin,opsin_constitution,identityorunverified(55756fccb); the best-effort tier gives a correct name or abstains (574514b0d).An atom-coverage check rejects names that drop atoms (0dfbc6eaa), and a grammar check tests brackets, hyphens and stereo placement before a name is emitted (317684d48).
Numbering ties are settled by the CIP criteria (P-14.4 (j)); the same input string always gives the same name; the centres jar is fixed to its tagged 1.2.1 release (866f91477).
Unsupported inputs, such as wildcard atoms, many inorganic compounds, unnameable salt parts and ring systems the engine cannot number, are declined with a reason (722efbc63).
Performance and robustness#
OPSIN runs in the same process through JPype instead of a new JVM for each name, which made naming about 2.6 times faster (502c251cd).
Memoization scoped to each call or molecule, with byte-identical output (342a40ab8); SMARTS patterns compiled once and ring lookups hash-bucketed (504084230).
A per-molecule work budget abstains on macrocycles that would take too long (90b265b39); large peptides, glycopeptides and oligosaccharides name within the time limit.
Command line, Python API and packaging#
Python API:
name_compound(), theOrthonymclass andname_with_tree(), which returns the name with its structured name tree (89441dced, b4247b9ee).The
orthonymcommand names single SMILES or batch files, with--emit-tier,--provenance(JSON) and--dump-tree(8dccfa14c); an abstention prints a clean line instead of ‘None’.orthonym --fetch-jarsdownloads the OPSIN and centres jars and checks each against its pinned SHA-256 checksum; the jars’ licences are listed in NOTICE.The package imports cleanly on a minimal install, declares lxml as a runtime dependency, and raises ValueError on an unparseable SMILES (29120be67).
Python 3.10+ and a Java 11+ runtime; MIT licence; public CI and release automation with release-please and PyPI Trusted Publishing (ea1ec9532).
Documentation#
A documentation site, https://steinbeck-lab.github.io/Orthonym/, covers install, a first name, tiers, how each name is checked, declines, accuracy, the Python API and the command line (6b3cb017c).
The README has a round-trip diagram, real example output, a quick start and credits (IUPAC 2013 recommendations, RDKit, OPSIN, centres), with guide pages on how it works, declines and accuracy (e4f4ae382).
Public docstrings and
--helptext describe each option in plain language; changelog, contributing, code of conduct and security files are included (a275e9958).The data record of the accompanying paper is at https://doi.org/10.5281/zenodo.22946586.
Development timeline#
January 2026: first public commits: chains, monocycles, benzene derivatives, Hantzsch-Widman heterocycles, fused and bridged ring systems, suffix seniority, CIP stereo, and the
orthonymcommand with a Python API.February 2026: ions, salts, zwitterions and radicals; von Baeyer polycycles, lactones, lactams and macrocycles; steroids, alkaloids, sugars and peptides; splitting of large molecules; a confidence score for every name.
March 2026: one shared substituent-naming path, parent selection by the P-44 criteria, fused names from components, dispiro names and splitting at several bonds.
April 2026: retained names, amino acids and sugars extended with OPSIN data and round-trip checks; a 300-molecule CIP validation set; one shared candidate pool.
May 2026: an OPSIN grammar check before emission, metallocenes, seleno and telluro acids, and
name_with_tree()with--dump-tree.June 2026: rules only (the experimental machine-learning fallback was removed); OPSIN parse and self-consistency checks on every name; centres as the default CIP source; indicated hydrogen; fusion numbering; carbohydrates, lipids, nucleotides and organometallics.
July 2026: output tiers,
--emit-tierand--provenance; OPSIN in the same process; isotope descriptors; phanes; stereo completeness; sigma-coordination names.August 2026: the best-effort tier gated on a full round trip; sulfinamides, spiro and fused breadth, semipolar oxides, salts, capped peptides, non-reducing oligosaccharides and the coordination-name table.
September 2026: every shown name passes its own round trip; radical names; systematic names for N-substituted amino acids; faster naming with scoped memoization;
--fetch-jarswith checksums; the documentation site; release 1.0.0 on 29 September.
Changes since the first publish#
Features#
more preferred IUPAC names, fewer lost names, two wrong-name classes closed, honest labels (e3a9a63)
nomenclature coverage and correctness updates (931fe63)
nomenclature coverage and correctness updates (1d835f6)
nomenclature coverage and correctness updates (327e28c)
radical names, safer salt names, stable numbering and Blue Book PIN spelling fixes (34ef0d3)
systematic names for N-substituted amino acids, long alkanes named again, and many preferred-name fixes (68f50d1)
Bug Fixes#
Documentation#
credits back in the README under Built on (2153b79)
lighter README, with guide pages for how it works, declines and accuracy (e4f4ae3)
project links in the README, the citation file, the package metadata and the documentation site (d2f3118)
the documentation site shows the Orthonym mark and the web app’s credit, and the README links the docs and the web app (7fc9231)
the Orthonym documentation site, with API docstrings and command-line help in plain language (6b3cb01)