# Orthonym documentation (full text) Version 1.0.3. Generated from the Markdown sources of the site. ```{toctree} :hidden: :caption: Start start/install start/first-name ``` ```{toctree} :hidden: :caption: Use use/command-line use/python use/batch use/provenance use/browser ``` ```{toctree} :hidden: :caption: Understand checking tiers declines accuracy how-it-works ``` ```{toctree} :hidden: :caption: Reference reference/python-api reference/command-line reference/internals/index ``` ```{toctree} :hidden: :caption: Project project/contributing project/changelog project/cite project/licence project/llms ``` # Install Orthonym is a Python package. It needs two things on your machine and downloads two more. | You need | Why | |:--|:--| | Python 3.10 or newer | the engine is written in Python | | A Java runtime, version 11 or newer, on your `PATH` | OPSIN, which reads names back for the round-trip check, and centres, which assigns the CIP stereodescriptors, are Java programs | Check the Java runtime with `java -version`. Any Java 11+ runtime works (for example OpenJDK or Temurin). ## Install the package ```console $ pip install "git+https://github.com/Steinbeck-Lab/Orthonym.git" $ orthonym --fetch-jars ``` The first line installs Orthonym and the packages it needs, among them RDKit, which reads the structures. The second line downloads the two Java programs Orthonym uses, checks each against the SHA-256 checksum recorded in Orthonym, and prints where they are: ```console $ orthonym --fetch-jars [orthonym] opsin 2.9.0: /home/you/.cache/orthonym/jars/opsin-cli-2.9.0-jar-with-dependencies.jar [orthonym] centres 1.2.1: /home/you/.cache/orthonym/jars/centres-cli-1.2.1.jar ``` `pip install` already tries this download for you. Run `orthonym --fetch-jars` once anyway: it also re-checks the files, and it tells you at once if something is missing. ## The two jars Orthonym does not ship any Java program. It uses two, each pinned to one version, and downloads them from their official releases: | Jar | Version | Used for | Licence | |:--|:--|:--|:--| | OPSIN, `opsin-cli-2.9.0-jar-with-dependencies.jar` | 2.9.0 | round-trip validation of names | MIT (the jar bundles jna-inchi, LGPL-2.1, and others) | | centres, `centres-cli-1.2.1.jar` | 1.2.1 | CIP stereodescriptors (*R*/*S*, *E*/*Z*) | BSD-2-Clause (the jar bundles CDK, LGPL-2.1+) | A missing jar is downloaded: when `pip` builds the package, on first use (unless `ORTHONYM_NO_DOWNLOAD=1` is set), or by `orthonym --fetch-jars`. Downloaded jars and jars in the jar directory are checked against the pinned SHA-256; a jar you point to with `ORTHONYM_OPSIN_JAR` or `ORTHONYM_CENTRES_JAR` is used as given. Orthonym stops with an error only when a jar can be neither found nor downloaded. Without a Java runtime it declines every molecule. ## Three settings for the jars | Setting | Effect | |:--|:--| | `ORTHONYM_OPSIN_JAR=/path/to/opsin-cli-2.9.0-jar-with-dependencies.jar` | use this OPSIN jar | | `ORTHONYM_CENTRES_JAR=/path/to/centres-cli-1.2.1.jar` | use this centres jar | | `ORTHONYM_JAR_DIR=/path/to/dir` | keep the downloaded jars here (the default is your user cache, `~/.cache/orthonym/jars`) | ## On a machine with no internet 1. On a machine with internet, run `orthonym --fetch-jars` and copy the two files it prints. 2. On the offline machine, point Orthonym at the copies: ```console $ export ORTHONYM_OPSIN_JAR=/opt/jars/opsin-cli-2.9.0-jar-with-dependencies.jar $ export ORTHONYM_CENTRES_JAR=/opt/jars/centres-cli-1.2.1.jar $ export ORTHONYM_NO_DOWNLOAD=1 ``` `ORTHONYM_NO_DOWNLOAD=1` stops Orthonym from trying to download anything. ## Naming without the jars `ORTHONYM_ALLOW_REDUCED=1` lets Orthonym name without the jars. Then no name is read back by OPSIN and the stereodescriptors come from RDKit instead of centres. Use it only when you know you want that. Next: [your first name](first-name.md). # Your first name Give Orthonym a structure written as SMILES. It prints the name. ## From the command line ```console $ orthonym "CCO" ethanol ``` Put the SMILES in quotes: characters such as `(`, `=` and `#` mean something to the shell. `python -m orthonym "CCO"` does the same. Five more, each the engine's own output: ```console $ orthonym "CC(=O)Oc1ccccc1C(=O)O" 2-(acetyloxy)benzoic acid $ orthonym "Cn1cnc2c1c(=O)n(C)c(=O)n2C" 1,3,7-trimethyl-3,7-dihydro-1H-purine-2,6-dione $ orthonym "C/C=C/C" (2E)-but-2-ene $ orthonym "C[C@H](O)CC" (2S)-butan-2-ol $ orthonym "CC(C)Cc1ccc(cc1)[C@@H](C)C(=O)O" (2R)-2-[4-(2-methylpropyl)phenyl]propanoic acid ``` That is aspirin, caffeine, *trans*-but-2-ene, (*S*)-butan-2-ol and (*R*)-ibuprofen. The stereodescriptors in the SMILES (`/`, `\`, `@`, `@@`) come back as *E*, *Z*, *R* and *S* in the name. ## From Python ```pycon >>> from orthonym import name_compound >>> name_compound("CCO") 'ethanol' >>> name_compound("CC(=O)Oc1ccccc1C(=O)O") '2-(acetyloxy)benzoic acid' >>> name_compound("Cn1cnc2c1c(=O)n(C)c(=O)n2C") '1,3,7-trimethyl-3,7-dihydro-1H-purine-2,6-dione' >>> name_compound("C/C=C/C") '(2E)-but-2-ene' >>> name_compound("C[C@H](O)CC") '(2S)-butan-2-ol' ``` ## How sure is the name? Ask for the provenance Add `--provenance` and Orthonym prints one row of JSON instead of the bare name. It says which tier the name earned and whether OPSIN read it back to your structure: ```console $ orthonym "CC(=O)Oc1ccccc1C(=O)O" --provenance | python -m json.tool { "name": "2-(acetyloxy)benzoic acid", "tier": "pin_verified", "is_pin": true, "source": "pin_path", "opsin": "verified", "gates_passed": [ "self_consistency" ], "gate_outcome": "self_consistency_verified", "formula": null, "limit_code": null, "stereo_unexpressed": false, "suffix_free_prefix_name": false, "prefix_order_fallback": false, "verified": "opsin" } ``` The two lines to read first: - `tier` is pin_verified: the strict path for the Preferred IUPAC Name built the name and verified it. The other tiers are on [Output tiers](../tiers.md). - `verified` is `opsin`: OPSIN read the name back to the same molecule. Every other field is explained on [The provenance row](../use/provenance.md). ## When there is no name Orthonym says so. The plain call prints a label in place of a name, and the provenance row gives the reason: ```console $ orthonym "O=[U](=O)=O" inorganic compound (not supported) ``` More on [declines](../declines.md). # Command line `orthonym` names one SMILES, or a file of them. `python -m orthonym` is the same program. Every option is listed below in plain words, grouped by what it is for; the [option reference](../reference/command-line.md) prints the program's own help text. ```console $ orthonym "c1ccccc1" benzene ``` ## Naming `orthonym "SMILES"` : Print the name of one structure. `--style {pin,general,cas}` : `pin` (the default) aims at the Preferred IUPAC Name. `general` allows a few general IUPAC forms where the recommendations offer one. `cas` is accepted and at present gives the same names as `pin`. `--enable-triviality-controller` : Where the IUPAC 2013 recommendations prefer a retained parent name (benzene, phenol, aniline, benzoic acid and others) to the systematic one, use it (P-15.1.8). Every change is checked by an OPSIN round trip. `ORTHONYM_ENABLE_TRIVIALITY_CONTROLLER=1` turns it on too. `--trivial` : When no preferred name can be built, also allow a retained trivial name that is not a preferred name. A preferred name that can be built is never replaced: glycerol stays `propane-1,2,3-triol`. Without this option a small table of retained trivial names is still used as a last resort at the wider tiers (the default tier declines those names); the provenance row labels them systematic_verified with source `trivial_retained`. ## Tiers `--emit-tier {pin,valid,complete,best-effort,full-coverage}` : Which names to return. The default, `pin`, returns a name only when the pipeline can build the preferred IUPAC name (PIN), that is when the strict PIN path built the name and verified it (pin_verified). The only exceptions are names from the natural-product and metal-complex lists, the name formats absent from OPSIN's grammar (for example inositols, phanes and thioperoxols), and PINs whose stereodescriptors OPSIN cannot read, for which the default tier compares the constitution. Otherwise it declines, with the reason code `NO_VERIFIED_PIN` when it built a name that is not a verified PIN. `valid` adds names from the general engine, `complete` adds general names for aromatic and heterocyclic ring systems, `best-effort` adds the last-resort producers, von Baeyer and spiro names for ring systems of up to 100 skeletal atoms and 11 rings (the other tiers build these names for ring systems of up to 40 skeletal atoms and 8 rings), and adducts with a one-atom ion such as chloride, and `full-coverage` adds the coordination-name builder for metal tetrapyrrole and corrin complexes. The wider tiers also return the names the default tier declines, each labelled with its tier. See [Output tiers](../tiers.md). ## Output `--provenance` : Print a JSON row instead of the bare name: the name, its tier, and how it was checked. One SMILES at a time. See [The provenance row](provenance.md). `--confidence` : Print the name with a coverage score, the part of the engine that built it, and the parts of the score. When no measurement was taken it says so: `Confidence: unverified (no coverage measurement was taken)`. `--verbose`, `-v` : Also print the SMILES and the style. With `--batch` and `--output`, also say how many lines were named and how many failed. `--dump-tree`, `--format {text,json}` : Print the parts of the name as a tree instead of the name. ```console $ orthonym --dump-tree "OC1CCCCC1" NameTree: cyclohexanol (class_id=general_acyclic, cite=P-14+P-23+P-44) +- parent_stem: 'cyclohex' |- suffix: 'ol' ``` ## Batch `--batch FILE`, `-b FILE` : Name every SMILES in the file, one per line. See [Batch files](batch.md). `--output FILE`, `-o FILE` : With `--batch`, write the lines to this file instead of the screen. ## Diagnostics These are for looking inside the engine. None of them changes a name, except `--binding-proof enforce`, which can decline one. `--fetch-jars` : Download the OPSIN and centres jars if they are missing, check them, print where they are, and stop. See [Install](../start/install.md). `--validation-stats` : After the name, print on the error stream how often the OPSIN grammar pre-check passed or repaired a candidate name. `--dispatch-stats` : After the name, print on the error stream which compound-class routes the engine took. `--engine-only` : Skip the strict path and print, as JSON, the general engine's own name, checked for atom coverage but not by OPSIN. Never a preferred name; do not use it as a name. `--binding-proof {off,audit,enforce}` : An extra check that every part of a general-engine name still maps onto its atoms in the final name. `audit` records the result and never changes the name; `enforce` also declines when the check fails. `--version`, `-V` : Print the installed version, in the form `orthonym 1.0.2`. ## Exit status `0` when the name (or the label for a decline) was printed; `1` for a SMILES that RDKit cannot read, or a batch with at least one such line; `2` when a jar can be neither found nor downloaded. # Python Everything the command line does is one import away. The full parameter lists are on the [Python API](../reference/python-api.md) page. ## One molecule, one name ```pycon >>> from orthonym import name_compound >>> name_compound("CC(C)Cc1ccc(cc1)[C@@H](C)C(=O)O") '(2R)-2-[4-(2-methylpropyl)phenyl]propanoic acid' ``` `name_compound` returns a string: the name, or a label when there is none. Tell the two apart with `is_failure_name`: ```pycon >>> from orthonym.errors import is_failure_name >>> is_failure_name(name_compound("O=[U](=O)=O")) True ``` A SMILES that RDKit cannot read raises `ValueError`: ```pycon >>> name_compound("C1CC") Traceback (most recent call last): ... ValueError: Invalid SMILES: C1CC ``` ## Many molecules, one set of options: `Orthonym` `name_compound` sets up a fresh engine for every call. For a loop, or for the wider tiers, build one `Orthonym` and reuse it. The keyword switches are the ones behind `--emit-tier`: | `--emit-tier` | `Orthonym(...)` | |:--|:--| | `pin` (default) | `Orthonym()` | | `valid` | `Orthonym(general_fallback=True)` | | `complete` | `Orthonym(general_fallback=True, allow_aromatic_general=True)` | | `best-effort` | `Orthonym(general_fallback=True, allow_aromatic_general=True, general_fallback_unverified=True)` | | `full-coverage` | as `best-effort`, plus `full_coverage=True` | ```pycon >>> from orthonym import Orthonym >>> namer = Orthonym(general_fallback=True) >>> namer.name("CCCCCCCCCCCCC/C=C/[C@H]([C@H](CO)N)O") '(2S,3R,4E)-2-aminooctadec-4-ene-1,3-diol' ``` ## The name and how it was checked: `name_tiered` `Orthonym.name_tiered` returns the provenance row as a dictionary, the same row `--provenance` prints: ```pycon >>> best_effort = Orthonym(general_fallback=True, allow_aromatic_general=True, ... general_fallback_unverified=True) >>> row = best_effort.name_tiered("C1CC[C@H]2CCCC[C@H]2C1") >>> row["name"], row["tier"], row["verified"] ('cis-bicyclo[4.4.0]decane', 'best_effort', 'opsin') ``` Each key is explained on [The provenance row](provenance.md). ## The parts of a name: `name_with_tree` ```pycon >>> from orthonym import name_with_tree >>> result = name_with_tree("OC1CCCCC1") >>> result.name 'cyclohexanol' >>> result.tree.parent_stem, result.tree.suffix ('cyclohex', 'ol') ``` ## Asking first: `classify_limit` `classify_limit` says whether a structure is out of scope, without raising: ```pycon >>> from orthonym import classify_limit >>> classify_limit("O=[U](=O)=O").code 'UNSUPPORTED_ELEMENT' >>> classify_limit("CCO") is None True ``` To get an exception instead of a label, pass `raise_on_limit=True`: ```pycon >>> from orthonym import OrthonymLimitError >>> try: ... Orthonym().name("O=[U](=O)=O", raise_on_limit=True) ... except OrthonymLimitError as err: ... print(err.code, "|", err.message) UNSUPPORTED_ELEMENT | inorganic compound (not supported) ``` # Batch files Put one SMILES per line in a text file. Empty lines are skipped. ```console $ cat molecules.smi CCO c1ccccc1 O=[U](=O)=O $ orthonym --batch molecules.smi CCO ethanol c1ccccc1 benzene O=[U](=O)=O inorganic compound (not supported) ``` Each output line is the SMILES, a tab, and the name or the label for a decline. Write the lines to a file with `--output`, and add `--verbose` for a count at the end: ```console $ orthonym --batch molecules.smi --output names.txt --verbose Processed 3 SMILES, 0 errors Output written to: names.txt ``` A line that RDKit cannot read does not stop the run. It gets `ERROR:` and the reason, and the exit status is then `1`: ```console $ orthonym --batch bad.smi CCO ethanol not-a-smiles ERROR: Invalid SMILES: not-a-smiles CC(=O)O acetic acid ``` Batch runs use the default tier. `--emit-tier` and `--provenance` do not apply to `--batch` yet. For tiers or provenance rows over many molecules, loop in Python with one `Orthonym` instance: ```python import json from orthonym import Orthonym namer = Orthonym(general_fallback=True) # the valid tier with open("molecules.smi") as src, open("rows.jsonl", "w") as out: for line in src: smiles = line.strip() if smiles: out.write(json.dumps({"smiles": smiles, **namer.name_tiered(smiles)}) + "\n") ``` # The provenance row `orthonym "SMILES" --provenance` prints one row of JSON instead of the bare name, and `Orthonym.name_tiered` returns the same row as a dictionary. It says what the name is, how it was built and how it was checked. ```console $ orthonym "Cn1cnc2c1c(=O)n(C)c(=O)n2C" --provenance {"name": "1,3,7-trimethyl-3,7-dihydro-1H-purine-2,6-dione", "tier": "pin_verified", "is_pin": true, "source": "pin_path", "opsin": "verified", "gates_passed": ["self_consistency"], "gate_outcome": "self_consistency_verified", "formula": null, "limit_code": null, "stereo_unexpressed": false, "suffix_free_prefix_name": false, "prefix_order_fallback": false, "verified": "opsin"} ``` ## Every field `name` : The name, or the label for a decline. With `--emit-tier` other than `pin`, a decline gives `null` here. `tier` : How the name was built: pin_verified, pin_unverified, systematic_verified, best_effort or abstain. See [Output tiers](../tiers.md). `is_pin` : `true` only for a certified Preferred IUPAC Name. `source` : Which part of the engine produced the name: `pin_path` (the strict path), `general_engine`, `trivial_retained` (the table of retained trivial names), `t4_floor` (a last-resort producer) or `abstain`. `opsin` : What the OPSIN check found: `verified`, `verified_constitution_only` (the constitution matched; the stereodescriptors were not compared), `unverified`, or `n/a` when there is no name. `gates_passed` : The checks this name passed, for example `self_consistency` (OPSIN read the name back to your structure), `atom_coverage` (every atom is named) or `full_key_round_trip`. `gate_outcome` : What the final OPSIN check did for this name: `self_consistency_verified`, `full_key_round_trip_verified`, `self_consistency_constitution_only`, `suppressed` (a candidate failed and was withdrawn), `not_run`, `unavailable` (no OPSIN check ran: the opt-in reduced mode without the jars), or `carveout:` for a name class that OPSIN cannot read. `formula` : The molecular formula, given when there is no name, so a decline still tells you what came in. `limit_code` : The reason code of a decline, for example `UNSUPPORTED_ELEMENT` or, at the default tier, `NO_VERIFIED_PIN` (the engine built a checked name that is not a verified PIN). See [Declines](../declines.md). `stereo_unexpressed` : `true` when a stereocentre of your structure is not stated in the name. `suffix_free_prefix_name` : `true` when the name states the principal characteristic group as a prefix with no suffix, a form the recommendations do not allow. `prefix_order_fallback` : `true` when a best_effort name cites its substituent prefixes in an order other than the alphanumerical one, because the name in the standard order did not pass its round trip for its stereodescriptors alone. Such a name is never a Preferred IUPAC Name; the default tier does not use it. `verified` : Whether the name was read back, in one word: `opsin` (OPSIN read the name back to the same molecule), `opsin_constitution` (OPSIN read it back with the same constitution; the stereodescriptors were not confirmed by OPSIN), `identity` (a name from an exact-match list: a metal-complex name found by your structure's exact InChIKey, or a natural-product parent name found by its exact structure; OPSIN cannot read these names), or `unverified` (no read-back recorded). ## Read `tier` and `verified` together The tier says how a name was built. The `verified` field says whether OPSIN read it back. The two can differ: at the default tier, *cis*-decalin gets a name in preferred-name form that OPSIN read back with the right constitution only. ```console $ orthonym "C1CC[C@H]2CCCC[C@H]2C1" --provenance {"name": "(4as,8as)-decahydronaphthalene", "tier": "pin_unverified", "is_pin": false, "source": "pin_path", "opsin": "verified_constitution_only", "gates_passed": ["self_consistency_constitution_only"], "gate_outcome": "self_consistency_constitution_only", "formula": null, "limit_code": null, "stereo_unexpressed": false, "suffix_free_prefix_name": false, "prefix_order_fallback": false, "verified": "opsin_constitution"} ``` `--provenance` works for one SMILES at a time; `--batch` ignores it. For many molecules, loop over `Orthonym.name_tiered` as shown on [Batch files](batch.md). # Use it in the browser The Orthonym web app, at [orthonym.decimer.ai](https://orthonym.decimer.ai), runs the same engine for chemists who do not want to install anything. Paste, upload or draw a structure, or send a whole file as one job and download the results. The app is a separate project, [Orthonym-Web on GitHub](https://github.com/Steinbeck-Lab/Orthonym-Web). ## What it shows for each structure Next to every result the app draws the structure OPSIN read back from the name, beside your input, so you can see the round trip yourself. Each result lands on one of five states. The mark carries the state by its shape, so the five stay distinct in greyscale: pin : The Preferred IUPAC Name. The engine's tier pin_verified: the strict PIN path built the name and verified it (OPSIN read it back to your structure; a name from the natural-product and metal-complex lists is matched to your exact structure instead). fallback : A checked name whose preferred status is not certified. The engine's tiers pin_unverified and systematic_verified. Most of these names were read back by OPSIN to your structure; at the default tier a name that OPSIN read back with its constitution only, or could not read, lands here too. The app's own read-back verdict under the name says which. best_effort : A name from a last-resort producer, not from the strict rules or the general engine. The engine's tier best_effort. The app prints its own read-back verdict under the name. abstain : The engine declined rather than guess, so no name was made. The engine's tier abstain. Error : No name was made because RDKit could not read the input, or naming failed part-way, or a batch line ran out of time. ## When the app's own check did not run When a result has no read-back from the app itself (for example with the read-back switched off), the app says "Name not checked here: no round-trip result" instead of the tier. When the app's read-back gives a different structure, it says "Round trip here gave a different structure". # How every name is checked A name that reads well can still describe the wrong molecule. Orthonym does not trust its own name: before a name leaves the engine, the name is checked against your structure, and a candidate that fails is withdrawn. ## The round trip 1. **You give a structure.** RDKit reads the SMILES. 2. **Orthonym writes a name** from the rules of the IUPAC 2013 recommendations. 3. **OPSIN reads the name back** into a structure. OPSIN is an independent program that turns names into structures; it never saw your SMILES. 4. **The two structures are compared** by InChIKey, a short code computed from a structure that changes with its constitution, charge and stereochemistry. When they match, the name ships. Here is caffeine, from the engine's own output: ## The three checks **1. Every atom is named.** Every heavy atom of your structure must be claimed by exactly one part of the name, and every part must appear in the final name. This check decides for the names of the general engine (the `valid`, `complete` and `best-effort` tiers). On the strict path of the default tier it is recorded but does not decide; the OPSIN checks below do. **2. OPSIN must be able to read the name.** At the `valid`, `complete` and `best-effort` tiers, a name OPSIN cannot read is not emitted. **3. What OPSIN read must be your molecule.** At the `valid`, `complete` and `best-effort` tiers the full InChIKey must match: constitution, charge and stereochemistry. A name with a missing or wrong stereodescriptor is therefore not emitted there. At the default tier the check accepts the same full InChIKey, or the same constitution, charge and protonation state with no stereodescriptor that disagrees with your structure. A candidate that fails is withdrawn, and the engine tries the next way of naming the molecule. When none passes, it declines and says why ([Declines](declines.md)). ## The one stereo repair When a name read back with the right constitution but without some of its stereodescriptors, the engine adds the missing descriptors from your structure once and has OPSIN read the new name. The name ships only if the full InChIKey then matches. The provenance row shows it as `gate_outcome` `stereo_omission_full_key_recomposed`. ## Where OPSIN cannot read the name Some correct names use words OPSIN does not know. Two kinds of name leave the engine without a full read-back, and the provenance row says so every time: - **Names from exact-match lists.** Metal tetrapyrrole complexes (hemes, chlorophylls, cobalamins, siroheme, coenzyme F430) and a table of retained natural-product parent names are matched to your structure exactly, by InChIKey or canonical SMILES, and the listed name is given. OPSIN cannot read these names at all. A list name is returned at every tier. It shows `verified` `identity` and the tier of the path that returned it: the natural-product parent names, and at the default tier the metal-complex names, are pin_verified. A natural-product parent name shows `gate_outcome` `carveout:np_stereoparent`. - **A few name classes at the default tier.** Some preferred-name forms (inositols, phanes and some anhydrides among them) are built by their rules and shipped at the default tier although OPSIN cannot read them, or reads only their constitution. The row shows `gate_outcome` `carveout:` or `self_consistency_constitution_only`, and the tier is lowered accordingly. At the `valid`, `complete` and `best-effort` tiers these names are declined instead. If OPSIN itself does not answer (the Java runtime times out or stops), the candidate is withdrawn and the row shows `gate_outcome` `suppressed`. A name ships with `gate_outcome` `unavailable` only in the opt-in reduced mode without the jars (`ORTHONYM_ALLOW_REDUCED=1`, [Install](start/install.md)). ## See it for yourself `--provenance` prints what happened to every name. [The provenance row](use/provenance.md) explains each field, and the web app draws the read-back structure next to yours ([In the browser](use/browser.md)). # Output tiers Every name Orthonym returns carries a tier. The tier says **how the name was built**. The `verified` field of the [provenance row](use/provenance.md) says **whether OPSIN read it back**. Read the two together. ## The five tiers pin_verified : The strict PIN path built the name and verified it: OPSIN read it back to your structure, or, for a name from the exact-match lists, your structure matched the list entry exactly: a metal-complex name by your structure's InChIKey, a natural-product parent name by its exact structure (`verified` `identity`). A compound without carbon (`sodium chloride`, `sulfuric acid`, `ammonia`) carries the tier its naming path gives it. `is_pin` is `true` only here. pin_unverified : A name in PIN form that only a breadth producer built, so its preferred status is not certified. It is also the tier of a strict-path name that OPSIN read back only in part (its constitution) or not at all. The `verified` field tells which. systematic_verified : A correct systematic name that is not the PIN: from the general engine (for example, a von Baeyer name for a fused ring system), from the table of retained trivial names, from the strict path when the name contains a part the engine records as not the preferred form (for example `(trimethylazaniumyl)acetate`, whose PIN is `(N,N-dimethylmethanaminiumyl)acetate`), or from the strict path for a class the Blue Book gives no PIN, such as organometallic compounds of the Group 1-12 metals (`ethenylsodium`) and compounds of aluminium, gallium, indium and thallium. best_effort : A name from the last-resort producers of the `best-effort` tier, or a name whose own string no round trip confirmed. abstain : No name. The row gives the reason code instead ([Declines](declines.md)). The marks are the ones the [web app](use/browser.md) uses. It shows the five tiers with four marks, because pin_unverified and systematic_verified share the FALLBACK mark, and it has a fifth state, Error, for input it cannot read. The shape carries the tier and the colour only agrees, so the marks stay distinct in greyscale: pin fallback best_effort abstain ## Choosing a tier: `--emit-tier` The default is `pin`. Wider tiers are opt-in: | `--emit-tier` | What it adds | |:--|:--| | `pin` *(default)* | A name only when the pipeline can build the preferred IUPAC name (PIN), that is when the strict PIN path built the name and verified it (pin_verified). The only exceptions are names from the natural-product and metal-complex lists, the name formats absent from OPSIN's grammar (for example inositols, phanes and thioperoxols), and PINs whose stereodescriptors OPSIN cannot read, for which the default tier compares the constitution. Otherwise it declines, with the reason code `NO_VERIFIED_PIN` when it built a name that is not a verified PIN. | | `valid` | Names from the general engine, each with an atom-coverage certificate and a full-InChIKey round trip. | | `complete` | General names for aromatic and heterocyclic ring systems as well. | | `best-effort` | The last-resort producers as well, and von Baeyer and spiro names for ring systems of up to 100 skeletal atoms and 11 rings (the other tiers build these names for ring systems of up to 40 skeletal atoms and 8 rings), and adducts with a one-atom ion such as chloride. | | `full-coverage` | The coordination-name builder for metal tetrapyrrole and corrin complexes as well; it builds a name or declines. | The wider tiers are not simply "the default plus more". They also return the names the default tier declines, each labelled with its tier. At `valid`, `complete` and `best-effort` every name must pass a full-InChIKey round trip, except a name from the natural-product and metal-complex lists (`verified` `identity`), so the name formats absent from OPSIN's grammar, which the default tier ships without a full read-back, are declined there ([How every name is checked](checking.md)). ## One molecule, three tiers *cis*-Decalin shows all of this at once. The default tier gives a name in preferred-name form that OPSIN read back with the right constitution only: ```console $ orthonym "C1CC[C@H]2CCCC[C@H]2C1" --provenance {"name": "(4as,8as)-decahydronaphthalene", "tier": "pin_unverified", "is_pin": false, "source": "pin_path", "opsin": "verified_constitution_only", "gates_passed": ["self_consistency_constitution_only"], "gate_outcome": "self_consistency_constitution_only", "formula": null, "limit_code": null, "stereo_unexpressed": false, "suffix_free_prefix_name": false, "prefix_order_fallback": false, "verified": "opsin_constitution"} ``` At `valid` that name does not pass the full round trip, and no general-engine name does either, so the engine declines: ```console $ orthonym "C1CC[C@H]2CCCC[C@H]2C1" --emit-tier valid (no name — UNNAMEABLE) ``` At `best-effort` a last-resort producer builds a name that OPSIN reads back to the same molecule, stereochemistry included: ```console $ orthonym "C1CC[C@H]2CCCC[C@H]2C1" --emit-tier best-effort --provenance {"name": "cis-bicyclo[4.4.0]decane", "tier": "best_effort", "is_pin": false, "source": "t4_floor", "opsin": "verified", "gates_passed": ["full_key_round_trip"], "gate_outcome": "full_key_round_trip_verified", "formula": null, "limit_code": null, "stereo_unexpressed": false, "suffix_free_prefix_name": false, "prefix_order_fallback": false, "verified": "opsin"} ``` At a tier other than `pin`, a decline prints `(no name — CODE)` on the command line, and the provenance row has `"name": null`. ## In Python The tiers are keyword switches of `Orthonym`; the table is on the [Python](use/python.md) page. # Declines When Orthonym cannot name a molecule, it says so. The plain call returns a label in place of a name, such as `inorganic compound (not supported)`, and the [provenance row](use/provenance.md) marks the row abstain with a reason code in `limit_code`. ```console $ orthonym "O=[U](=O)=O" --provenance | python -m json.tool OPSIN validity gate suppressed unparseable name: 'unknown' { "name": "inorganic compound (not supported)", "tier": "abstain", "is_pin": false, "source": "abstain", "opsin": "n/a", "gates_passed": [], "gate_outcome": "suppressed", "formula": "O3U", "limit_code": "UNSUPPORTED_ELEMENT", "stereo_unexpressed": false, "suffix_free_prefix_name": false, "prefix_order_fallback": false, "verified": "unverified" } ``` The first line is a log message on the error stream: the engine withdrew a candidate that OPSIN could not read. The JSON row is on the normal output. ## The reason codes | `limit_code` | Label in place of the name | When | |:--|:--|:--| | `UNSUPPORTED_ELEMENT` | `inorganic compound (not supported)`, or ` compound (not supported)` | an element outside the ones the rules cover | | `WILDCARD_ATOMS` | `compound with wildcard atoms (not supported)` | the structure has a `*` atom | | `STRUCTURE_TOO_LARGE` | `unknown organic compound` | an unnamed structure with more than 125 heavy atoms: a label given after the attempt, since the engine sets no size limit and attempts every input | | `ISOLATED_ATOM` | `unknown organic compound` | a single heavy atom | | `UNSUPPORTED_RING_SYSTEM` | `unknown organic compound` | a ring system the engine recognises but cannot name correctly yet (with `raise_on_limit=True`) | | `UNNAMEABLE` | `unknown organic compound` | an organic structure no candidate name passed the checks for | | `NO_VERIFIED_PIN` | `unknown organic compound` | the default tier built and checked a name, but not a verified PIN and not one of its exceptions; `--emit-tier best-effort` returns that name with its tier | Most organic declines carry the general `UNNAMEABLE` (no candidate name passed the checks) or, at the default tier, `NO_VERIFIED_PIN` (a checked name exists, but not the PIN). Betaine shows the second case. Its strict-path name contains a part the engine records as not the preferred form, so the default tier declines, and a wider tier returns the name labelled systematic_verified: ```console $ orthonym "C[N+](C)(C)CC(=O)[O-]" --provenance {"name": "unknown organic compound", "tier": "abstain", "is_pin": false, "source": "abstain", "opsin": "n/a", "gates_passed": [], "gate_outcome": "suppressed", "formula": "C5H11NO2", "limit_code": "NO_VERIFIED_PIN", "stereo_unexpressed": false, "suffix_free_prefix_name": false, "prefix_order_fallback": false, "verified": "unverified"} $ orthonym "C[N+](C)(C)CC(=O)[O-]" --emit-tier valid (trimethylazaniumyl)acetate ``` At a tier other than `pin`, the plain command line prints `(no name — CODE)` instead of the label: ```console $ orthonym --emit-tier best-effort "O=[U](=O)=O" (no name — UNSUPPORTED_ELEMENT) ``` ## In code `orthonym.errors.is_failure_name` tells a label from a name: ```pycon >>> from orthonym import name_compound >>> from orthonym.errors import is_failure_name >>> is_failure_name(name_compound("O=[U](=O)=O")) True >>> is_failure_name(name_compound("CCO")) False ``` `orthonym.classify_limit` returns the reason without raising, and `raise_on_limit=True` turns a decline into an `OrthonymLimitError` ([Python](use/python.md)). For a molecule the default tier declines with `NO_VERIFIED_PIN`, `classify_limit` returns that code and `raise_on_limit=True` raises `OrthonymLimitError` with it. A decline is never silent: a molecule that gets no name gets a label, and `--provenance` gives a reason code. `--batch` prints the label only. # Accuracy Orthonym is judged by what it gets **exactly right** and by what it gets **wrong**. ## The three measures **Round-trip exact match.** Name the structure, have OPSIN read the name back, and compare the full InChIKey of what OPSIN read with the input's. Every molecule counts. A declined molecule counts as a miss, and so does a name OPSIN cannot read. **Wrong structures.** An emitted name that OPSIN reads back to a different constitution. The design target is zero. Names that OPSIN cannot read are reported separately, in the column "Unreadable name"; they are never counted as right. **Preferred-name conformance.** Names compared, character for character apart from spacing and superscript digits, with the worked examples that the IUPAC 2013 recommendations mark as the preferred name (PIN) or the preselected name and that OPSIN can turn into a structure, measured with the `best-effort` tier switched on. ## Measured results Orthonym 1.0.0 at the `best-effort` tier, every name read back by OPSIN 2.9.0 with its radical option and compared by full InChIKey. | Set | Molecules | Named (round-trip exact) | Declined | Unreadable name | Other layers | |:--|--:|--:|--:|--:|--:| | Held-out ChEBI 2,000 | 2,000 | 1,991 (99.55 %) | 9 | 0 | 0 | | QM9 | 133,885 | 133,860 (99.98 %) | 25 | 0 | 0 | | ChEBI | 111,843 | 107,686 (96.28 %) | 4,118 | 39 | 0 | | PubChem 500,000 | 500,000 | 499,171 (99.83 %) | 829 | 0 | 0 | | ZINC22 500,000 | 500,000 | 488,407 (97.68 %) | 11,593 | 0 | 0 | | PubChem 1,000,000 | 1,000,000 | 957,523 (95.75 %) | 42,468 | 1 | 8 | The columns other than Molecules add up to Molecules. Named (round-trip exact): the name read back to the input's full InChIKey; the percentage is of Molecules. Declined: no name, with a reason. Unreadable name: an emitted name OPSIN could not read; all are names from the engine's exact-match lists (metal tetrapyrrole complexes and a natural-product parent), which OPSIN does not know. Other layers: the same constitution with a different protonation state. Wrong structures, a name that reads back to a different constitution: none on any set. The comparison with other generators and the full analysis are in the paper. ## Where the numbers come from The counts are the Orthonym rows of the run records deposited with the paper's data (`run_records/RESCORE_SUMMARY.md`). The held-out ChEBI set is a split of ChEBI structures that no development run had read. The site build recomputes every percentage in the table from the counts in `guide/_data/accuracy.json`, checks that each row adds up, and stops if the table disagrees. # How it works A SMILES string goes in. The engine reads the structure, applies the nomenclature rules, assembles a name and, for nearly every name, checks it with an OPSIN round trip before it leaves ([How every name is checked](checking.md)). A few name classes that OPSIN cannot read leave on their construction alone; `--provenance` marks them. Four stages do the work. The orchestrator, `namer.py`, runs them and holds the final OPSIN check. | Stage | What it does | |:--|:--| | **Perception** | RDKit reads the structure; Orthonym finds rings, characteristic groups and stereocentres. The CIP descriptors come from the centres labeller; RDKit's CIP labeller fills the few double bonds centres leaves unlabelled and takes over for a molecule centres cannot label. | | **Rules** | Seniority, the parent hydride, locants, alphanumerical order and spelling, built as nomenclature classes from the IUPAC 2013 recommendations. Names the recommendations list one by one (retained and natural-product names), the names of metal tetrapyrrole complexes (ChEBI names, matched by exact InChIKey) and a last-resort table of trivial names are looked up by exact structure. | | **Assembly** | Chains, rings, Hantzsch–Widman heterocycles, fused, bridged (von Baeyer) and spiro systems, and the characteristic-group families: acids, esters, amides, amines, nitriles and more. | | **Validation** | The OPSIN round trip that decides whether a name leaves the engine, with named exceptions for classes OPSIN cannot read, and, for names from the general engine, an atom-coverage certificate. | Compound classes are tried in a fixed order; a class that cannot build a name declines, and the next one is tried. There is no special case for a single molecule and no rewriting of an emitted name. ## The source tree ```text src/orthonym/ ├── namer.py # orchestration and the final OPSIN check ├── cli.py # the command line ├── jars.py # finding, downloading and checking the OPSIN and centres jars ├── perception/ # structure perception: rings, characteristic groups, CIP stereo ├── routing/ # choosing the compound class ├── decomposition/ # naming large structures from their fragments ├── rules/ # IUPAC nomenclature rules ├── assembly/ # name assembly: locants, ordering, selection, the general engine ├── validation/ # OPSIN round-trip helpers and atom-coverage checks ├── metrics/ # provenance, tier labels and reason codes └── data/ # naming tables ``` Every module has a page under [Internals](reference/internals/index.md). ## Built on RDKit, OPSIN and centres, on the IUPAC 2013 recommendations. The credits and references are in the README, under [Built on](https://github.com/Steinbeck-Lab/Orthonym#built-on). # Python API The public names of the `orthonym` package. Everything else is [internal](internals/index.md); the counters of `Orthonym` (`get_validation_stats`, `get_dispatch_stats` and the others) are documented there, under `orthonym.namer`. | Name | What it is for | |:--|:--| | [`name_compound`](#orthonym.name_compound) | One molecule, one name | | [`Orthonym`](#orthonym.Orthonym) | All options, one engine for many molecules | | [`Orthonym.name_tiered`](#orthonym.Orthonym.name_tiered) | The provenance row: the name, its tier, and every field of how it was checked | | [`name_with_tree`](#orthonym.name_with_tree) | The name and the tree of its parts | | [`NamingResult`](#orthonym.NamingResult), [`NameTreeNode`](#orthonym.NameTreeNode) | What `name_with_tree` returns | | [`classify_limit`](#orthonym.classify_limit), [`OrthonymLimitError`](#orthonym.OrthonymLimitError) | Why a structure is out of scope | | [`is_failure_name`](#orthonym.errors.is_failure_name) | Tell a label from a name | Every example below is the engine's own output. ## Naming ## The provenance fields ## The parts of a name ## Declines ## Version `orthonym.__version__` is the installed version, a string such as `'1.0.2'`. ## Python API docstrings ### orthonym.name_compound(smiles: str, style: str = 'pin', include_confidence: bool = False, *, enable_triviality_controller: bool = False, enable_group_splitting: bool = False, trivial_fallback: bool = False, general_fallback: Optional[bool] = None, general_fallback_unverified: Optional[bool] = None, allow_aromatic_general: Optional[bool] = None, full_coverage: Optional[bool] = None, raise_on_limit: bool = False, binding_proof: str = 'off') Name one molecule. Reads a structure written as SMILES and returns its IUPAC name as a string. At the default settings a name is returned only when the strict path for the Preferred IUPAC Name built it and verified it, or when it is one of the default tier's exceptions (the exact-match list names, the few name formats OPSIN cannot read, a PIN whose stereodescriptors OPSIN cannot read); otherwise, and where no name passes, it is a label such as ``'unknown organic compound'`` or ``'inorganic compound (not supported)'``. The wider tiers (``general_fallback`` and the options below) also return names that are not the preferred name. Use :func:`orthonym.errors.is_failure_name` to tell a label from a name, and :meth:`Orthonym.name_tiered` to learn which tier a name earned. Parameters ---------- smiles: str The structure, as a SMILES string. style: {"pin", "general", "cas"}, default "pin" Naming style. ``"pin"`` aims at the Preferred IUPAC Name. ``"general"`` allows a few general IUPAC forms where the recommendations offer one (for example some adduct names and axial stereodescriptors). ``"cas"`` is accepted and at present gives the same names as ``"pin"``. include_confidence: bool, default False Return a dictionary instead of a string. Its ``"name"`` key holds the name, ``"handler"`` the part of the engine that built it, and ``"confidence"`` and ``"factors"`` a coverage score and its parts. ``"confidence"`` is ``None`` when no measurement was taken. enable_triviality_controller: bool, default False Where the recommendations prefer a retained parent name to the systematic one, use the retained name. Each change is checked by an OPSIN round trip. The environment setting ``ORTHONYM_ENABLE_TRIVIALITY_CONTROLLER=1`` turns this on as well. enable_group_splitting: bool, default False When an ester or thioester group inside a larger molecule has no prefix form, write it as its parts (``oxo`` plus ``ethoxy``, for example) instead of giving up. Each split name must pass an OPSIN round trip. The environment setting ``ORTHONYM_ENABLE_GROUP_SPLITTING=1`` turns this on as well. trivial_fallback: bool, default False When no preferred name can be built, also allow a retained trivial name that is not a preferred name. A molecule whose preferred name can be built keeps it. Turns on ``enable_triviality_controller``. general_fallback: bool or None, default None When the strict path declines, try the general naming engine (the ``valid`` tier of the command line). Its names must pass an atom-coverage certificate and a full-InChIKey round trip. ``None`` means "off" for a call you make; the engine uses it to pass the setting on when it names parts of a molecule. general_fallback_unverified: bool or None, default None Also switch on the ``best-effort`` tier: the last-resort producers, von Baeyer and spiro names for ring systems of up to 100 skeletal atoms and 11 rings (the other tiers build these names for ring systems of up to 40 skeletal atoms and 8 rings), and adducts with a one-atom ion. Their names still have to pass the full-InChIKey round trip. ``None`` as for ``general_fallback``. allow_aromatic_general: bool or None, default None Let the general engine also name aromatic and heterocyclic ring systems (the ``complete`` tier). ``None`` as for ``general_fallback``. full_coverage: bool or None, default None Also try the coordination-name builder for metal tetrapyrrole and corrin complexes (the ``full-coverage`` tier). It builds a name or declines. ``None`` as for ``general_fallback``. raise_on_limit: bool, default False Raise:class:`OrthonymLimitError` for a structure the engine cannot handle, instead of returning a label. binding_proof: {"off", "audit", "enforce"}, default "off" An extra check that every part of a general-engine name still maps onto the atoms it names in the final string. ``"audit"`` records the result and never changes the name; ``"enforce"`` also declines when the check fails. Returns ------- str or dict The name, or a label that says why no name was given. A dictionary when ``include_confidence`` is true. Raises ------ ValueError If RDKit cannot read the SMILES, or ``binding_proof`` is not one of the three values. OrthonymLimitError If ``raise_on_limit`` is true and the structure is out of scope. Examples -------- >>> from orthonym import name_compound >>> name_compound("CC(C)Cc1ccc(cc1)[C@@H](C)C(=O)O") '(2R)-2-[4-(2-methylpropyl)phenyl]propanoic acid' >>> name_compound("O=[U](=O)=O") 'inorganic compound (not supported)' ### orthonym.name_with_tree(smiles: str, style: str = 'pin') Name one molecule and return the parts of the name. A shortcut for ``Orthonym(style=style).name_with_tree(smiles)``. Parameters ---------- smiles: str The structure, as a SMILES string. style: {"pin", "general", "cas"}, default "pin" Naming style, as for:func:`name_compound`. Returns ------- NamingResult The name (the same string:func:`name_compound` returns), the tree of its parts as a:class:`NameTreeNode`, and a map from atom index to locant where one was recorded. Raises ------ ValueError If RDKit cannot read the SMILES. Examples -------- >>> from orthonym import name_with_tree >>> result = name_with_tree("OC1CCCCC1") >>> result.name 'cyclohexanol' >>> result.tree.parent_stem 'cyclohex' >>> result.tree.suffix 'ol' ### orthonym.Orthonym The naming engine, with every option. Build one instance and name as many molecules with it as you like; the options apply to every call.:func:`name_compound` builds a fresh instance for each call. Creating an instance checks that the OPSIN and centres jars are present and stops with an error if they are not (run ``orthonym --fetch-jars``). Set ``ORTHONYM_ALLOW_REDUCED=1`` to name without them, with no OPSIN check. Parameters ---------- style: {"pin", "general", "cas"}, default "pin" Naming style, as for:func:`name_compound`. enable_triviality_controller, enable_group_splitting, trivial_fallback: bool As for:func:`name_compound`. All default to False. general_fallback, general_fallback_unverified, allow_aromatic_general, full_coverage: bool The switches behind the command line's ``--emit-tier``. All default to False, which is the default tier. ``valid`` sets ``general_fallback``; ``complete`` adds ``allow_aromatic_general``; ``best-effort`` adds ``general_fallback_unverified``; ``full-coverage`` adds ``full_coverage``. binding_proof: {"off", "audit", "enforce"}, default "off" As for:func:`name_compound`. Notes ----- The parameters whose names start with an underscore are for the engine's own use and tests; leave them at their defaults. Examples -------- >>> from orthonym import Orthonym >>> namer = Orthonym >>> namer.name("CCO") 'ethanol' >>> namer.name("CC(=O)O") 'acetic acid' ### orthonym.Orthonym.name(self, smiles: str, *, raise_on_limit: bool = False) -> str Name one molecule with this instance's options. Parameters ---------- smiles: str The structure, as a SMILES string. raise_on_limit: bool, default False Raise:class:`OrthonymLimitError` for a structure the engine cannot handle, instead of returning a label. Returns ------- str The name, or a label that says why no name was given (see :func:`orthonym.errors.is_failure_name`). Raises ------ ValueError If RDKit cannot read the SMILES. OrthonymLimitError If ``raise_on_limit`` is true and the structure is out of scope. For a ring system the engine cannot name yet the code is ``UNSUPPORTED_RING_SYSTEM``; when that happens inside a part of the molecule, the code reported can be the more general ``UNNAMEABLE``. Examples -------- >>> from orthonym import Orthonym >>> Orthonym.name("C/C=C/C") '(2E)-but-2-ene' ### orthonym.Orthonym.name_tiered(self, smiles: str) -> dict Name one molecule and say how the name was made and checked. This is the row the command line prints with ``--provenance``. Parameters ---------- smiles: str The structure, as a SMILES string. Returns ------- dict ``name`` The name, or a label when there is none. ``tier`` How the name was built: ``pin_verified`` (the strict path for the Preferred IUPAC Name built it, certified it and OPSIN read it back), ``pin_unverified`` (a name in preferred-name form that OPSIN read back, but whose preferred status is not certified: a producer outside the strict path built it or a part of it), ``systematic_verified`` (a checked systematic name that is not the preferred name, from the general engine, a table of retained names, or the strict path when the name contains a part the engine records as not the preferred form), ``best_effort`` (the last-resort producers, or a name whose own string no round trip confirmed) or ``abstain`` (no name). ``is_pin`` True only for a certified Preferred IUPAC Name. ``source`` Which part of the engine produced the name, for example ``pin_path``, ``general_engine``, ``trivial_retained`` or ``abstain``. ``opsin`` What the OPSIN check found: ``verified``, ``verified_constitution_only``, ``unverified`` or ``n/a``. ``gates_passed`` The checks this name passed, for example ``self_consistency`` (OPSIN read the name back to your structure) and ``atom_coverage``. ``gate_outcome`` What the final OPSIN check did for this name, for example ``self_consistency_verified``, ``suppressed`` or ``not_run``, or ``carveout:`` for a class OPSIN cannot read. ``formula`` The molecular formula, given when there is no name. ``limit_code`` The reason code when there is no name, for example ``UNSUPPORTED_ELEMENT``, or ``NO_VERIFIED_PIN`` when the default tier built a name that is not a verified PIN (a wider tier returns it). ``stereo_unexpressed`` True when a stereocentre of the input is not stated in the name. ``suffix_free_prefix_name`` True when the name states the principal characteristic group as a prefix with no suffix. ``prefix_order_fallback`` True when the name cites its substituent prefixes out of the alphanumerical order because the round trip of the ordered spelling fails at the stereo layer only (a stereocentre read differently by the SMILES hand-off, not by the name); such a name is best_effort, never a PIN. ``verified`` ``opsin`` (OPSIN read the whole name back to the same molecule), ``opsin_constitution`` (read back with the same constitution; the stereodescriptors were not confirmed by OPSIN), ``identity`` (a name from an exact-match list: a metal-complex name found by the input's exact InChIKey, or a natural-product parent name found by its exact structure; OPSIN cannot read these names) or ``unverified`` (no read-back recorded). Examples -------- >>> from orthonym import Orthonym >>> row = Orthonym.name_tiered("Cn1cnc2c1c(=O)n(C)c(=O)n2C") >>> row["name"], row["tier"], row["verified"] ('1,3,7-trimethyl-3,7-dihydro-1H-purine-2,6-dione', 'pin_verified', 'opsin') ### orthonym.Orthonym.name_with_tree(self, smiles: str) Name one molecule and return the parts of the name. Parameters ---------- smiles: str The structure, as a SMILES string. Returns ------- NamingResult ``name`` is the same string:meth:`name` returns; ``tree`` is the :class:`NameTreeNode` of its parts (a single coarse node when the part of the engine that built the name records no finer structure); ``atom_to_locant_hint`` maps atom indices to locants where one was recorded, else ``None``. Raises ------ ValueError If RDKit cannot read the SMILES. Examples -------- >>> from orthonym import Orthonym >>> Orthonym.name_with_tree("OC1CCCCC1").name 'cyclohexanol' ### orthonym.Orthonym.name_with_confidence(self, smiles: str) -> dict Name one molecule and return a coverage score with it. Returns ------- dict ``name`` (the name), ``confidence`` (a score from 0 to 1, or ``None`` when no measurement was taken), ``verification``, ``factors`` (the parts of the score), ``handler`` (the part of the engine that built the name), and further keys for the atom-to-locant map and the reason for a decline (``limit``, ``abstention``). Raises ------ ValueError If RDKit cannot read the SMILES. Examples -------- >>> from orthonym import Orthonym >>> Orthonym.name_with_confidence("CCO")["name"] 'ethanol' ### orthonym.NamingResult A name together with the tree of its parts. Returned by:func:`orthonym.name_with_tree` and :meth:`orthonym.Orthonym.name_with_tree`. It is a named tuple of three fields. Attributes ---------- name: str The name, the same string:func:`orthonym.name_compound` returns. tree: NameTreeNode or None The parts of the name. atom_to_locant_hint: dict of int to int or None Atom index to locant, where the part of the engine that built the name recorded it. Examples -------- >>> from orthonym import name_with_tree >>> name, tree, hint = name_with_tree("OC1CCCCC1") >>> name 'cyclohexanol' ### orthonym.NameTreeNode One part of a name, and the parts inside it. A name is written in the order stereodescriptors, prefixes (in alphanumerical order), parent, indicated hydrogen, unsaturation and suffix (IUPAC. A node holds those pieces for one parent; each prefix is a node of its own, so a substituent with its own substituents is a subtree. The node cannot be changed after it is made. Attributes ---------- parent_stem: str The parent, for example ``'cyclohex'``. locants: tuple of int Locants of the suffix. suffix: str or None The suffix, for example ``'ol'``. prefixes: tuple of NameTreeNode The substituent prefixes, each a node. stereo: str or None The stereodescriptor part, for example ``'(2R)'``. indicated_h: tuple of int Locants of indicated hydrogen. unsaturation_locants: tuple of (tuple of int, tuple of int) Locants of double and of triple bonds. class_id: str The compound class that built the node. multiplicative_prefix: str or None A multiplying prefix such as ``'di'``. parenthesization_hint: bool True when the prefix must be written in enclosing marks. iupac_section_cite: str or None The section of the recommendations the node follows, for example ``''``. fragment_legacy: object or None An older representation of the same part, kept for the engine's own use. Examples -------- >>> from orthonym import name_with_tree >>> tree = name_with_tree("OC1CCCCC1").tree >>> tree.parent_stem, tree.suffix ('cyclohex', 'ol') ### orthonym.classify_limit(smiles: str) -> Optional[orthonym.errors.OrthonymLimitError] Say whether a structure is out of scope, without raising. A wildcard atom (``*``) is refused at once. Any other structure is named with the default settings; if that gives a label instead of a name, the reason is returned. Parameters ---------- smiles: str The structure, as a SMILES string. Returns ------- OrthonymLimitError or None The reason, with its code (for example ``UNSUPPORTED_ELEMENT``, ``WILDCARD_ATOMS``, or ``NO_VERIFIED_PIN`` when the default settings built a name that is not a verified preferred IUPAC name) and message. ``None`` when the structure gets a name, and also when RDKit cannot read the SMILES at all (that is a reading error, not a scope limit). Examples -------- >>> from orthonym import classify_limit >>> classify_limit("O=[U](=O)=O").code 'UNSUPPORTED_ELEMENT' >>> classify_limit("CC*").code 'WILDCARD_ATOMS' >>> classify_limit("CCO") is None True ### orthonym.OrthonymLimitError The reason a structure is out of scope. Returned by:func:`orthonym.classify_limit`, and raised when you ask for it with ``raise_on_limit=True``. It tells "cannot handle this structure" apart from a name. Attributes ---------- code: str The reason code: ``WILDCARD_ATOMS``, ``UNSUPPORTED_ELEMENT``, ``ISOLATED_ATOM``, ``STRUCTURE_TOO_LARGE``, ``UNSUPPORTED_RING_SYSTEM``, ``UNNAMEABLE`` or ``NO_VERIFIED_PIN`` (the default tier built a name but not a verified preferred IUPAC name; a wider tier returns it). message: str The label the plain call returns in place of a name. design_note_ref: str or None A reference to the design note the code follows. smiles: str or None The input, when known. Examples -------- >>> from orthonym import Orthonym, OrthonymLimitError >>> try: ... Orthonym.name("O=[U](=O)=O", raise_on_limit=True) ... except OrthonymLimitError as err: ... print(err.code, "|", err.message) UNSUPPORTED_ELEMENT | inorganic compound (not supported) ### orthonym.errors.is_failure_name(name: Optional[str]) -> bool Tell a label from a name. When the engine cannot name a structure, the plain call returns a label in place of a name, such as ``'inorganic compound (not supported)'`` or ``'unknown organic compound'``. This function is true for those labels and for an empty result, and false for a real name. Parameters ---------- name: str or None What a naming call returned. Returns ------- bool True when ``name`` is empty or a label (it contains ``unknown`` or ``(not supported)``). Examples -------- >>> from orthonym import name_compound >>> from orthonym.errors import is_failure_name >>> is_failure_name(name_compound("O=[U](=O)=O")) True >>> is_failure_name(name_compound("CCO")) False # Command-line options The program's own help text, read from its option parser when the site is built. The same options in plain words, grouped by what they are for, are on [Command line](../use/command-line.md). # Internals ```{note} :class: ot-internal Internal API. Names and behaviour may change between releases. ``` Every module of the engine has a page here, generated from its docstrings when the site is built. The docstrings are notes written by and for the developers; for using Orthonym, the [Python API](../python-api.md) is the page to read. The modules are grouped by the stage of the engine they belong to ([How it works](../../how-it-works.md)). ```{toctree} :maxdepth: 1 generated/perception generated/rules generated/assembly generated/validation generated/data generated/decomposition generated/metrics generated/routing generated/top-level ``` # Contributing to Orthonym Thanks for your interest in improving Orthonym. This guide covers how to set up a development environment, run the tests, and contribute a change. ## Development setup ```bash git clone https://github.com/Steinbeck-Lab/Orthonym.git cd Orthonym python -m venv .venv && source .venv/bin/activate pip install -e ".[dev]" ``` A **Java runtime (JRE 11+)** must be on your `PATH`: Orthonym validates candidate names by round-tripping them through OPSIN, which is a Java program. The OPSIN and centres jars are not part of the repository; `pip install` tries to fetch them when it builds the package, a missing jar is downloaded on first use, and `orthonym --fetch-jars` fetches them or re-checks the ones in the jar directory at any time (see the README, "The OPSIN and centres jars"). ## Running tests Run tests on targeted file sets (the OPSIN-backed tests need a Java runtime and the jars): ```bash python -m pytest tests/unit/rules/test_multiplicative.py -q python -m pytest tests/unit/rules/test_d1_coordination_v36.py -q ``` Markers (`unit`, `integration`, `roundtrip`, `slow`, `benchmark`) are defined in `pyproject.toml`. ## How Orthonym is built The engine perceives structure (`perception/`), dispatches by compound class (`routing/`, `decomposition/`), applies nomenclature rules (`rules/`), and assembles the name (`assembly/`), drawing on naming tables in `data/`. The orchestrator `namer.py` runs the final OPSIN check, using helpers in `validation/` (which also holds the atom-coverage check); `metrics/` records the provenance and the tier of each name. [`guide/how-it-works.md`](guide/how-it-works.md) describes the main stages. ``` Orthonym/ ├── src/orthonym/ │ ├── namer.py # the naming pipeline and the final OPSIN check │ ├── cli.py # the command line │ ├── perception/ # structure perception: rings, characteristic groups, CIP stereo │ ├── routing/ # compound-class dispatch │ ├── decomposition/ # fragment-based naming of large structures │ ├── rules/ # IUPAC nomenclature rules │ ├── assembly/ # name assembly: locants, ordering, selection │ ├── validation/ # OPSIN round-trip and atom-coverage checks │ ├── metrics/ # provenance, tiers and abstention codes │ └── data/ # naming tables └── tests/ # unit and integration tests ``` To add a compound class, add a test that pins the expected name and cites the governing IUPAC rule. Release notes are in [`CHANGELOG.md`](CHANGELOG.md), vulnerability reports go by [`SECURITY.md`](SECURITY.md), and questions and bug reports go to [open an issue](https://github.com/Steinbeck-Lab/Orthonym/issues/new/choose). ## Contribution guidelines 1. **Fix at the root cause.** Fix a naming defect at the rule or data source that produced it — not with a per-molecule special case or a blind rewrite of the emitted name. 2. **Never emit a wrong structure.** A change must not cause Orthonym to emit a name that describes a different molecule than the input. When a preferred name cannot be built with confidence, degrade to a correct systematic name rather than guess. 3. **Cite the rule.** When a change touches nomenclature correctness, add or extend a test that pins the expected name, and cite the governing IUPAC rule in the code or test. 4. **Keep it deterministic.** Any algorithm that resolves a choice (ring numbering, locant assignment, ordering) must be deterministic. ## Submitting a change 1. Fork the repository and create a topic branch. 2. Make your change with tests. 3. Run the relevant tests and confirm they pass. 4. Open a pull request describing the change and the IUPAC rule it implements or corrects. ## Code of conduct This project follows the [Contributor Covenant](CODE_OF_CONDUCT.md). By participating, you agree to uphold it. ## License By contributing, you agree that your contributions are licensed under the MIT License. # Changelog All notable changes to Orthonym are documented here. The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and this project adheres to [Semantic Versioning](https://semver.org/). ## [1.0.3](https://github.com/Steinbeck-Lab/Orthonym/compare/v1.0.2...v1.0.3) (2026-10-01) ### Bug Fixes * bridged fused PINs on naphthalene and anthracene parents, larger ring systems at the best-effort tier, more PIN classes and exact spellings ([7bdb0d2](https://github.com/Steinbeck-Lab/Orthonym/commit/7bdb0d2e6c2f6817610f5d9a4607a2202e6e716e)) ## [1.0.2](https://github.com/Steinbeck-Lab/Orthonym/compare/v1.0.1...v1.0.2) (2026-09-30) ### Bug Fixes * engine fixes and optimizations, tier updates ([6cab387](https://github.com/Steinbeck-Lab/Orthonym/commit/6cab38789a73993310344acf4d15dfd882391fde)) ## [1.0.1](https://github.com/Steinbeck-Lab/Orthonym/compare/v1.0.0...v1.0.1) (2026-09-29) The naming engine is the same as in 1.0.0: the Python code differs only in comments and docstrings. The PyPI package now lists all three authors and Development Status 5 - Production/Stable, and the documentation cites the Zenodo archive. ### Documentation * the Zenodo DOI in the README, the docs and CITATION.cff ([9415508](https://github.com/Steinbeck-Lab/Orthonym/commit/94155089a8a4ed7180bd7f060c9f66a9859cb968)) ## [1.0.0](https://github.com/Steinbeck-Lab/Orthonym/releases/tag/v1.0.0) (2026-09-29) Orthonym 1.0.0 is the first public release of a deterministic, rule-based generator that turns a SMILES string into an IUPAC name, aiming at the Preferred IUPAC Name (PIN) of the IUPAC 2013 recommendations. Each name is parsed back by OPSIN and must match the input structure by full InChIKey before it is shown; otherwise Orthonym declines and gives the reason, and the few names OPSIN cannot read in full are labelled as such. The release gathers 2,799 public commits made from late January to 29 September 2026, which took the engine from chains, monocycles and the main functional groups to fused, bridged, spiro and phane ring systems, stereodescriptors, charged species, isotopes, natural products, peptides, glycans, lipids and metal compounds. On the corpora of the accompanying paper, run with the paper's recipe on the release branch, it gave no wrong names. ### Highlights - Rules only, with no neural network and no sampling: the same SMILES always gives the same name. - Names are read back by OPSIN and compared with the input in constitution, charge and stereo; a name that fails is withdrawn, and when none is left Orthonym declines with a reason. - A PIN tier for names built and certified on the strict preferred-name path, and opt-in wider tiers up to best effort (`--emit-tier`) whose names must pass a full-InChIKey round trip (metal-complex list names excepted). - `--provenance` prints each result as JSON with its tier, its source and a `verified` field that says how the name was checked. - 0 wrong names on the paper's corpora (measured on the release branch with the paper's recipe); round-trip exact: QM9 133,860 of 133,885, ChEBI 107,946 of 111,843, PubChem 500k 499,666 of 500,000, ZINC22 500k 488,428 of 500,000. - Ring systems: monocycles, Hantzsch-Widman heterocycles, fused systems with IUPAC numbering, von Baeyer polycycles, spiro compounds, phanes and ring assemblies. - Stereo: CIP R/S from the centres engine, E/Z, ring cis/trans and pseudoasymmetric r/s, each checked against the input's CIP labels. - Ions, salts, zwitterions, radicals, isotope-labelled compounds, natural products, peptides, oligosaccharides, lipids, nucleotides and organometallics. - OPSIN runs in the same process through JPype; the OPSIN and centres jars are downloaded at install and checked against pinned SHA-256 checksums. - Documentation at https://steinbeck-lab.github.io/Orthonym/ and a web app at https://orthonym.decimer.ai (source: https://github.com/Steinbeck-Lab/Orthonym-Web). ### Nomenclature coverage **Chains and substituents** - Acyclic chains with IUPAC numerical terms, lowest locants and double- and triple-bond locants ([956329481](https://github.com/Steinbeck-Lab/Orthonym/commit/956329481)), including unbranched alkanes of 80 carbons and more ([68f50d128](https://github.com/Steinbeck-Lab/Orthonym/commit/0941b15ba)). - One substituent-naming path shared by the ester, lactone, lactam, amide, acid-halide, anhydride and benzene namers, so substituents are not silently dropped ([08b6e0d62](https://github.com/Steinbeck-Lab/Orthonym/commit/08b6e0d62)); alkenyl and alkynyl prefixes ([d63e482b3](https://github.com/Steinbeck-Lab/Orthonym/commit/d63e482b3)). - Nested enclosing marks for compound substituents, including N-substituents on amides (P-16.5) ([eb5d40c69](https://github.com/Steinbeck-Lab/Orthonym/commit/eb5d40c69)); alphabetical tie-breaking of chain orientation ([ec90679f5](https://github.com/Steinbeck-Lab/Orthonym/commit/ec90679f5)). - Multiplicative names (bis/tris/tetrakis) that keep every identical arm, with bridges such as sulfanediyl, peroxy and disulfanediyl ([bcea5da52](https://github.com/Steinbeck-Lab/Orthonym/commit/22899d493), [5b2fc9c87](https://github.com/Steinbeck-Lab/Orthonym/commit/8040c91f0)). **Parent selection and functional groups** - The parent chain or ring is chosen by the P-44 criteria in order: principal groups, ring or chain, length, multiple bonds, lowest locants ([863ff10db](https://github.com/Steinbeck-Lab/Orthonym/commit/863ff10db), [6cf607ac7](https://github.com/Steinbeck-Lab/Orthonym/commit/60cb38e0d)). - Suffix seniority follows the functional-group order of P-41, and lower-ranked groups are cited as prefixes ([4b6158840](https://github.com/Steinbeck-Lab/Orthonym/commit/4b6158840), [d396a18c9](https://github.com/Steinbeck-Lab/Orthonym/commit/d396a18c9)). - Nitriles, amides, esters, acid halides, anhydrides, hydroperoxides, carbamic and thiocarboxylic acids, and functional class names for oximes, N-oxides, isocyanates and carbamates ([b73e049cf](https://github.com/Steinbeck-Lab/Orthonym/commit/b73e049cf)). - Seleno and telluro acids ([24fb4c019](https://github.com/Steinbeck-Lab/Orthonym/commit/57703b4a7)), sulfinamides ([62bd3f515](https://github.com/Steinbeck-Lab/Orthonym/commit/85b4775db)), amidines, imidates, carbamoylamino and ylidene prefixes, and principal chains of As, Sb, Se and Te oxoacids. **Rings: monocycles, Hantzsch-Widman and fused systems** - Cycloalkanes, cycloalkenes and Hantzsch-Widman names for 3- to 10-membered heterocycles, such as oxolane and thiazolidine ([785a1e0c3](https://github.com/Steinbeck-Lab/Orthonym/commit/785a1e0c3)); unsaturated 7- and 8-membered heterocycles take Hantzsch-Widman names ([c6599a59b](https://github.com/Steinbeck-Lab/Orthonym/commit/15fea9ec3)). - Retained fused heterocycles with IUPAC peripheral numbering ([87146be5e](https://github.com/Steinbeck-Lab/Orthonym/commit/87146be5e)); fused systems not in the dictionary are named from their components, the base component chosen by P-25.3.1.3 ([a821b6b53](https://github.com/Steinbeck-Lab/Orthonym/commit/a821b6b53)). - A fusion-numbering engine for cata-fused arenes and mixed 5/6-membered systems, which declines ring systems it does not cover ([6a82e08bf](https://github.com/Steinbeck-Lab/Orthonym/commit/a6c417359)). - Substituted two-component ortho-fused heterocycles and naphtho-fused systems ([bb07d4ba6](https://github.com/Steinbeck-Lab/Orthonym/commit/e5f747895), [a2a12971b](https://github.com/Steinbeck-Lab/Orthonym/commit/829491b05)). - Indicated hydrogen, added indicated hydrogen and hydro prefixes, with cyclic ketones and quinones named as PIN diones ([7472e8dca](https://github.com/Steinbeck-Lab/Orthonym/commit/b83a56176)); chromones and coumarins on the 1-benzopyran parent ([757586d58](https://github.com/Steinbeck-Lab/Orthonym/commit/d475b7d3b)). **Rings: von Baeyer, spiro, phanes and ring assemblies** - Von Baeyer names with secondary and zero-atom bridges and heteroatom replacement ([5fde4e994](https://github.com/Steinbeck-Lab/Orthonym/commit/5fde4e994)); main ring and main bridge chosen for PIN descriptors, preferring the largest main bridge ([ce8a95d7b](https://github.com/Steinbeck-Lab/Orthonym/commit/ec665fee3)). - Spiro names, including dispiro and heterospiro systems, spiro compounds of von Baeyer and fluorene components ([e991c19c7](https://github.com/Steinbeck-Lab/Orthonym/commit/97939fc53)), and dispiro and trispiro systems of fused components ([8e66cc181](https://github.com/Steinbeck-Lab/Orthonym/commit/73a5820fe)). - Preferred names for monocyclic all-benzene homophanes ([a3c878a2a](https://github.com/Steinbeck-Lab/Orthonym/commit/8f71aee6b)); ring assemblies with primes (1,1':4',1''-terphenyl), enclosing marks and multipliers up to deci. **Stereochemistry** - R/S and E/Z with locants ([6bb4693a9](https://github.com/Steinbeck-Lab/Orthonym/commit/6bb4693a9)), with the centres engine as the default source of CIP labels ([0537e1517](https://github.com/Steinbeck-Lab/Orthonym/commit/a0b546eb7)). - Descriptors on substituents, oximes, esters, lactones, lactams, fused heterocycles and polycycles, with CIP labels recomputed on each fragment ([646fe501a](https://github.com/Steinbeck-Lab/Orthonym/commit/646fe501a), [0179bca74](https://github.com/Steinbeck-Lab/Orthonym/commit/0179bca74)). - Chain and ring stereocentres expressed in full; a name that leaves out or invents a stereocentre is rejected ([c242eb2c7](https://github.com/Steinbeck-Lab/Orthonym/commit/90455d8ad)); cis/trans on rings with two stereocentres ([9288fd0e3](https://github.com/Steinbeck-Lab/Orthonym/commit/b92fcdc82)). - Steroid alpha/beta descriptors ([e4056e902](https://github.com/Steinbeck-Lab/Orthonym/commit/5807751d6)). **Charged species, salts, zwitterions and radicals** - Each input is classified as neutral, ion, zwitterion, salt or radical ([867cc6026](https://github.com/Steinbeck-Lab/Orthonym/commit/867cc6026)), and salts are named as cation plus anion ([512aed128](https://github.com/Steinbeck-Lab/Orthonym/commit/5bb88bb06)). - Semipolar oxides (oxido/-ium), azide, diazo, nitrooxy and isocyano groups ([2803b3a97](https://github.com/Steinbeck-Lab/Orthonym/commit/70e39c1fa)); nitro and N-oxide groups are named as neutral groups ([37461389e](https://github.com/Steinbeck-Lab/Orthonym/commit/37461389e)). - Stereo ionic salts, protonated diamines, hemisalts and mixed multi-anion salts; a salt with a part that cannot be named is refused ([2c622fff5](https://github.com/Steinbeck-Lab/Orthonym/commit/d78da5586)). - Carbon, chalcogen, ring-N, N-oxyl, acyl, nitrene and multicentre radicals, each shown only after a round trip that compares the radical sites ([34ef0d3af](https://github.com/Steinbeck-Lab/Orthonym/commit/f575d9083)). **Isotopes** - Isotopic descriptors ordered by symbol and mass, with counts as subscripts, as in (1,1-2H2) ([50331d9a9](https://github.com/Steinbeck-Lab/Orthonym/commit/7b3beaaa9)); labelled retained names fall back to a systematic parent, and labels go on the right word of a salt name ([f5eaa886a](https://github.com/Steinbeck-Lab/Orthonym/commit/e6c977a10)). **Natural products and retained names** - Steroid and alkaloid scaffolds named from the scaffold, with IUPAC atom numbering for nine steroid skeletons ([55ed43bb3](https://github.com/Steinbeck-Lab/Orthonym/commit/55ed43bb3)); nor-, homo- and seco- modifications ([c1bc100a2](https://github.com/Steinbeck-Lab/Orthonym/commit/c1bc100a2)). - Terpene stereoparents (IUPAC Table 10.1c), a complete D/L set of aldonic acids and xylitol ([7b1175c72](https://github.com/Steinbeck-Lab/Orthonym/commit/15ddd47cd)); ring stereo on saturated steroid scaffolds ([815c950f6](https://github.com/Steinbeck-Lab/Orthonym/commit/f7b9e25d8)). - Retained names that are not PINs give way to systematic ones (acetone becomes propan-2-one, picric acid 2,4,6-trinitrophenol); retained names from OPSIN data are round-trip checked before use ([163007b89](https://github.com/Steinbeck-Lab/Orthonym/commit/163007b89), [5a696fd61](https://github.com/Steinbeck-Lab/Orthonym/commit/5a696fd61)). **Peptides, glycans, lipids and nucleotides** - Peptides as acylamino chains, including capped termini, ester C-termini and non-standard backbones ([6515385de](https://github.com/Steinbeck-Lab/Orthonym/commit/6515385de), [d5371532a](https://github.com/Steinbeck-Lab/Orthonym/commit/ba7acaf0a)). - An amino acid with a substituent on its nitrogen takes its systematic name, e.g. (2S)-1-acetylpyrrolidine-2-carboxylic acid ([68f50d128](https://github.com/Steinbeck-Lab/Orthonym/commit/0941b15ba)). - About 50 retained sugar names and glycosides ([769783748](https://github.com/Steinbeck-Lab/Orthonym/commit/769783748)); disaccharides and oligosaccharides, including non-reducing ones such as raffinose ([419c678b8](https://github.com/Steinbeck-Lab/Orthonym/commit/63a6159c2)). - Triglycerides and polyol polyesters ([6a85190b4](https://github.com/Steinbeck-Lab/Orthonym/commit/6a85190b4)); phosphatidylcholine, phosphatidylethanolamine, phosphatidylserine, phosphatidic acid, ceramides and sphingolipids ([9db641d5f](https://github.com/Steinbeck-Lab/Orthonym/commit/4855a60ba)). - Nucleosides and nucleotides with modified sugars or substituted bases ([d64ccb421](https://github.com/Steinbeck-Lab/Orthonym/commit/131f21ab0)); acyl-CoA molecules on their 9H-purine parent ([9f236276e](https://github.com/Steinbeck-Lab/Orthonym/commit/67ed5a4db)). **Organometallics and inorganic parents** - Metallocenes such as ferrocene, mononuclear metal carbonyls and half-sandwich complexes ([9e0fe162a](https://github.com/Steinbeck-Lab/Orthonym/commit/863c1c025)); sigma-bonded main-group organometallics ([95ef27193](https://github.com/Steinbeck-Lab/Orthonym/commit/43d25fc74)). - Additive names for sigma-coordinated and metallacycle compounds; compounds with two metals are refused rather than guessed ([a147f99a1](https://github.com/Steinbeck-Lab/Orthonym/commit/45c9e2929)). - An exact-InChIKey table of retained coordination names drawn from ChEBI, including corrinoid precursors ([c75c88549](https://github.com/Steinbeck-Lab/Orthonym/commit/253ed2828)), and hydrate word forms (mono, di, hemi, sesqui). - Parent hydrides of Group 15, the chalcogens, the halogens, boron and Group 14, polyazanes and lambda-convention hydrides ([376445ebf](https://github.com/Steinbeck-Lab/Orthonym/commit/74c6d148b)). **Large molecules and mixtures** - Large molecules split at ester, amide, ether, glycosidic, thioester, phosphodiester and sulfonamide bonds, including several bonds of mixed types, with each piece named and the name rebuilt ([26c75704c](https://github.com/Steinbeck-Lab/Orthonym/commit/26c75704c), [c2f90e2d0](https://github.com/Steinbeck-Lab/Orthonym/commit/c2f90e2d0)). - Neutral multi-component SMILES, such as cocrystals, named one component at a time in a fixed order ([f43b8e3e4](https://github.com/Steinbeck-Lab/Orthonym/commit/f43b8e3e4)). ### Checks and correctness - Every public entry point checks its own names by an OPSIN round trip that compares constitution, charge and stereo ([0be43e0e5](https://github.com/Steinbeck-Lab/Orthonym/commit/5d3ad6079)); a name that OPSIN reads as a different molecule is not shown ([5f43cd2bf](https://github.com/Steinbeck-Lab/Orthonym/commit/360b1eb86)). - The default tier also names a few classes OPSIN cannot read in full: names from exact-match lists (metal-complex and natural-product parent names), a few name forms OPSIN's grammar lacks or misreads, and names whose stereodescriptors OPSIN cannot parse (constitution confirmed by OPSIN, each descriptor checked against its CIP label). `--provenance` marks each; the wider tiers ship none of them except the metal-complex list names. - Tiers reported by `--provenance`: `pin_verified`, `pin_unverified`, `systematic_verified`, `best_effort` and `abstain`; `verified` is `opsin`, `opsin_constitution`, `identity` or `unverified` ([55756fccb](https://github.com/Steinbeck-Lab/Orthonym/commit/04569243d)); the best-effort tier gives a correct name or abstains ([574514b0d](https://github.com/Steinbeck-Lab/Orthonym/commit/4bac21b81)). - An atom-coverage check rejects names that drop atoms ([0dfbc6eaa](https://github.com/Steinbeck-Lab/Orthonym/commit/0dfbc6eaa)), and a grammar check tests brackets, hyphens and stereo placement before a name is emitted ([317684d48](https://github.com/Steinbeck-Lab/Orthonym/commit/773403827)). - Numbering ties are settled by the CIP criteria (P-14.4 (j)); the same input string always gives the same name; the centres jar is fixed to its tagged 1.2.1 release ([866f91477](https://github.com/Steinbeck-Lab/Orthonym/commit/583619cbc)). - Unsupported inputs, such as wildcard atoms, many inorganic compounds, unnameable salt parts and ring systems the engine cannot number, are declined with a reason ([722efbc63](https://github.com/Steinbeck-Lab/Orthonym/commit/304972556)). ### Performance and robustness - OPSIN runs in the same process through JPype instead of a new JVM for each name, which made naming about 2.6 times faster ([502c251cd](https://github.com/Steinbeck-Lab/Orthonym/commit/7db7d39c5)). - Memoization scoped to each call or molecule, with byte-identical output ([342a40ab8](https://github.com/Steinbeck-Lab/Orthonym/commit/0cb1d833c)); SMARTS patterns compiled once and ring lookups hash-bucketed ([504084230](https://github.com/Steinbeck-Lab/Orthonym/commit/d64f9f0a2)). - A per-molecule work budget abstains on macrocycles that would take too long ([90b265b39](https://github.com/Steinbeck-Lab/Orthonym/commit/7a88ab14c)); large peptides, glycopeptides and oligosaccharides name within the time limit. ### Command line, Python API and packaging - Python API: `name_compound()`, the `Orthonym` class and `name_with_tree()`, which returns the name with its structured name tree ([89441dced](https://github.com/Steinbeck-Lab/Orthonym/commit/89441dced), [b4247b9ee](https://github.com/Steinbeck-Lab/Orthonym/commit/580eecd60)). - The `orthonym` command names single SMILES or batch files, with `--emit-tier`, `--provenance` (JSON) and `--dump-tree` ([8dccfa14c](https://github.com/Steinbeck-Lab/Orthonym/commit/115ec06c4)); an abstention prints a clean line instead of 'None'. - `orthonym --fetch-jars` downloads the OPSIN and centres jars and checks each against its pinned SHA-256 checksum; the jars' licences are listed in NOTICE. - The package imports cleanly on a minimal install, declares lxml as a runtime dependency, and raises ValueError on an unparseable SMILES ([29120be67](https://github.com/Steinbeck-Lab/Orthonym/commit/21b8953cd)). - Python 3.10+ and a Java 11+ runtime; MIT licence; public CI and release automation with release-please and PyPI Trusted Publishing ([ea1ec9532](https://github.com/Steinbeck-Lab/Orthonym/commit/dacac2da9)). ### Documentation - A documentation site, https://steinbeck-lab.github.io/Orthonym/, covers install, a first name, tiers, how each name is checked, declines, accuracy, the Python API and the command line ([6b3cb017c](https://github.com/Steinbeck-Lab/Orthonym/commit/b123e4ab3)). - The README has a round-trip diagram, real example output, a quick start and credits (IUPAC 2013 recommendations, RDKit, OPSIN, centres), with guide pages on how it works, declines and accuracy ([e4f4ae382](https://github.com/Steinbeck-Lab/Orthonym/commit/deb7cbdaf)). - Public docstrings and `--help` text describe each option in plain language; changelog, contributing, code of conduct and security files are included ([a275e9958](https://github.com/Steinbeck-Lab/Orthonym/commit/99a94f073)). - The data record of the accompanying paper is at https://doi.org/10.5281/zenodo.22946586. ### Development timeline - **January 2026:** first public commits: chains, monocycles, benzene derivatives, Hantzsch-Widman heterocycles, fused and bridged ring systems, suffix seniority, CIP stereo, and the `orthonym` command with a Python API. - **February 2026:** ions, salts, zwitterions and radicals; von Baeyer polycycles, lactones, lactams and macrocycles; steroids, alkaloids, sugars and peptides; splitting of large molecules; a confidence score for every name. - **March 2026:** one shared substituent-naming path, parent selection by the P-44 criteria, fused names from components, dispiro names and splitting at several bonds. - **April 2026:** retained names, amino acids and sugars extended with OPSIN data and round-trip checks; a 300-molecule CIP validation set; one shared candidate pool. - **May 2026:** an OPSIN grammar check before emission, metallocenes, seleno and telluro acids, and `name_with_tree()` with `--dump-tree`. - **June 2026:** rules only (the experimental machine-learning fallback was removed); OPSIN parse and self-consistency checks on every name; centres as the default CIP source; indicated hydrogen; fusion numbering; carbohydrates, lipids, nucleotides and organometallics. - **July 2026:** output tiers, `--emit-tier` and `--provenance`; OPSIN in the same process; isotope descriptors; phanes; stereo completeness; sigma-coordination names. - **August 2026:** the best-effort tier gated on a full round trip; sulfinamides, spiro and fused breadth, semipolar oxides, salts, capped peptides, non-reducing oligosaccharides and the coordination-name table. - **September 2026:** every shown name passes its own round trip; radical names; systematic names for N-substituted amino acids; faster naming with scoped memoization; `--fetch-jars` with checksums; the documentation site; release 1.0.0 on 29 September. ### Changes since the first publish #### Features * more preferred IUPAC names, fewer lost names, two wrong-name classes closed, honest labels ([e3a9a63](https://github.com/Steinbeck-Lab/Orthonym/commit/e56e9821bda098c3ea264aada0de7909fdac729d)) * nomenclature coverage and correctness updates ([931fe63](https://github.com/Steinbeck-Lab/Orthonym/commit/c2beb0bc6df1d8e4b81f179e6eb3488adf784903)) * nomenclature coverage and correctness updates ([1d835f6](https://github.com/Steinbeck-Lab/Orthonym/commit/d468f7b65a760f039086b1b8f7fc352537aa7e53)) * nomenclature coverage and correctness updates ([327e28c](https://github.com/Steinbeck-Lab/Orthonym/commit/44c23872a36ebc859181dc709f59919ad52feab4)) * radical names, safer salt names, stable numbering and Blue Book PIN spelling fixes ([34ef0d3](https://github.com/Steinbeck-Lab/Orthonym/commit/f575d9083cfa7402a9ddf5de217aac6be529d1cf)) * systematic names for N-substituted amino acids, long alkanes named again, and many preferred-name fixes ([68f50d1](https://github.com/Steinbeck-Lab/Orthonym/commit/0941b15ba3db53d7596e505b1b947dc5ac6c7c28)) #### Bug Fixes * every shown name passes its own round trip, N-substituted amino acids get systematic names, and ChEBI losses are restored ([0be43e0](https://github.com/Steinbeck-Lab/Orthonym/commit/5d3ad607909f2181afcf14b10ec7121c66331bc5)) * honest tier labels, faster naming of very large molecules, and every paper-named ChEBI structure named again ([eb25c33](https://github.com/Steinbeck-Lab/Orthonym/commit/748a6031fb0e32c0fdb101bdeff5cf90225f9531)) * nomenclature correctness updates ([57eaa92](https://github.com/Steinbeck-Lab/Orthonym/commit/7c5c6c52e75fc8c3af0b717a1d685cf520406f4a)) #### Documentation * credits back in the README under Built on ([2153b79](https://github.com/Steinbeck-Lab/Orthonym/commit/2e5e1c9830505b96b11e2876585dff4a4d3a4eca)) * lighter README, with guide pages for how it works, declines and accuracy ([e4f4ae3](https://github.com/Steinbeck-Lab/Orthonym/commit/deb7cbdaf8665f96ccb8859248c64df2eb28fbe9)) * project links in the README, the citation file, the package metadata and the documentation site ([d2f3118](https://github.com/Steinbeck-Lab/Orthonym/commit/4d1bfb9d10509589199efd3729de6880b60f351e)) * the documentation site shows the Orthonym mark and the web app's credit, and the README links the docs and the web app ([7fc9231](https://github.com/Steinbeck-Lab/Orthonym/commit/87099ae2f74fe35039111e956544ca9070ed8f7e)) * the Orthonym documentation site, with API docstrings and command-line help in plain language ([6b3cb01](https://github.com/Steinbeck-Lab/Orthonym/commit/b123e4ab3f7b0964f96e75510d42502114263a75)) # How to cite A paper describing Orthonym is in preparation. Until it is published, please cite the software. GitHub's **Cite this repository** button, built from [`CITATION.cff`](https://github.com/Steinbeck-Lab/Orthonym/blob/main/CITATION.cff), gives the entry in APA and BibTeX. In BibTeX: ```bibtex @software{orthonym, author = {Rajan, Kohulan and Zielesny, Achim and Steinbeck, Christoph}, title = {{Orthonym}}, version = {1.0.3}, year = {2026}, doi = {10.5281/zenodo.23044199}, url = {https://github.com/Steinbeck-Lab/Orthonym} } ``` Every release is archived on Zenodo. The DOI [10.5281/zenodo.23044199](https://doi.org/10.5281/zenodo.23044199) stands for all versions and resolves to the newest one. Orthonym is built on the IUPAC 2013 recommendations, OPSIN, centres and RDKit. When you describe how a name was checked, please cite them too; the references are in the README, under [Built on](https://github.com/Steinbeck-Lab/Orthonym#built-on). # Licence Orthonym is released under the MIT licence; the text is in [`LICENSE`](https://github.com/Steinbeck-Lab/Orthonym/blob/main/LICENSE). The OPSIN and centres jars are not part of Orthonym. They are downloaded from their official releases and keep their own licences: OPSIN is MIT (the jar bundles jna-inchi, LGPL-2.1, and others) and centres is BSD-2-Clause (the jar bundles CDK, LGPL-2.1+). The licences and sources of all third-party components are listed in [`NOTICE`](https://github.com/Steinbeck-Lab/Orthonym/blob/main/NOTICE). The typefaces of this site, Saira Condensed, Public Sans and JetBrains Mono, are under the SIL Open Font License and are served from the site itself. # llms.txt Two plain-text files at the root of this site are written for AI agents and other programs that read documentation: Both are generated from the Markdown sources of this site every time it is built, so they say what the pages say.