orthonym.rules.locants#
Note
Internal API. Names and behaviour may change between releases.
Locant assignment utilities for IUPAC nomenclature.
Locants map atom positions in the principal chain to 1-indexed IUPAC locant numbers. This module provides:
build_atom_to_locant: Convert atom indices to locant mapping
compare_locant_sets: First-point-of-difference comparison
orient_chain: Apply IUPAC 2013 chain orientation criteria
get_functional_group_locants: Resolve FG atoms to chain locants
- IUPAC 2013 chain orientation criteria (applied in order):
Lowest locants for principal characteristic group
Lowest locants for multiple bonds (as a set)
Lowest locants for double bonds (if tie with triple bonds)
Lowest locants for substituents (detachable prefixes)
Reference: IUPAC 2013 Blue Book,,,
- orthonym.rules.locants.element_seniority_rank(symbol)#
Numbering-seniority rank of an element symbol (lower = senior).
Unlisted symbols return a sentinel rank past every listed element so they sort last in a deterministic, total order. Source: (DD4).
- orthonym.rules.locants.build_atom_to_locant(principal_chain)#
Create a mapping from RDKit atom indices to 1-indexed IUPAC locants.
The locant numbering follows the order of the principal chain: the first atom in the chain gets locant 1, the second gets locant 2, etc.
- Parameters:
principal_chain (List[int]) – Ordered list of atom indices forming the principal chain. The ordering must already reflect the correct numbering direction (see orient_chain).
- Returns:
Dict mapping atom_idx -> locant (1-indexed). Empty dict if principal_chain is empty.
- Return type:
Dict[int, int]
Examples
>>> build_atom_to_locant([5, 3, 1, 0]) {5: 1, 3: 2, 1: 3, 0: 4} >>> build_atom_to_locant() {} >>> build_atom_to_locant([0]) {0: 1}
- orthonym.rules.locants.compare_locant_sets(set_a, set_b)#
Compare two locant sets using IUPAC first-point-of-difference rule.
This is NOT the sum-of-locants method. Both sets are sorted ascending, then compared term-by-term. The set with the lower value at the first point of difference is preferred.
If all compared elements are equal, the shorter set wins (fewer locants needed means simpler name). If completely identical, returns 0.
a phase extension (,): also accepts
List[Tuple[int, str]]with first-point-of-difference semantics for fusion atoms. When one list contains tuples and the other contains ints, all ints are coerced to(n, '')tuples internally — empty string sorts before any letter, preserving the IUPAC convention that locant4is “lower” than locant4a. Truly heterogeneous lists are rejected by_assert_homogeneous_locants(raisesValueError).- Parameters:
set_a (List[int | Tuple[int, str]]) – First locant set (unsorted or sorted). Elements are int or
(int, str)tuples; mixed allowed only if all ints can be coerced via the(n, '')padding.set_b (List[int | Tuple[int, str]]) – Second locant set (same type contract as set_a).
- Returns:
- -1 if set_a is preferred (lower at first difference)
0 if sets are equal 1 if set_b is preferred (lower at first difference)
- Return type:
int
Examples
>>> compare_locant_sets([2, 3, 5], [3, 4, 6]) -1 >>> compare_locant_sets([2, 4, 5], [2, 3, 5]) 1 >>> compare_locant_sets([2, 3], [2, 3]) 0 >>> compare_locant_sets([(4, ''), (5, '')], [(4, 'a'), (5, '')]) -1 >>> compare_locant_sets([(4, 'a'), (5, '')], [(4, 'b'), (5, '')]) -1
Source: https://iupac.qmul.ac.uk/BlueBook/P1.html, Source: a phase (locant type safety); a phase /
- orthonym.rules.locants.compound_aware_multibond_set(a2l, bonds)#
(1)/(2) [the Blue Book,:16683] multi-bond / double- bond-only locant-SET comparator, for use ONLY when a ring triple bond is present alongside ring double bond(s) (bi-/polycyclic von Baeyer systems with mixed unsaturation).
Unlike (2)’s explicit “any number in parentheses is ignored” (:16657 – a NAIVE cited-locant compare, i.e.
compare_locant_setsonsorted(min(a2l[a], a2l[b]) for a, b in bonds)), ‘s own worked example for this criterion (bicyclo[8.3.1]tetradeca-4,6,10- trien-2-yne, PIN,:16693) is a COUNTER-EXAMPLE to that naive reading: the REJECTED numbering (...-1(13),4,6-trien-8-yne) has a lower NAIVE cited-locant set ({1,4,6,8}<{2,4,6,10}) yet loses. Each bond is therefore ranked by(is_compound, cited_locant)before sorting, so a bond needing a compound locant never wins a rank position on the strength of its numerically-low cited half – “the seniority of single locants over compound locants… extended to von Baeyer systems” (BBv2 Note at:16639) applied as the tie-break within this SET.Returns a tuple of
(is_compound, cited_locant)pairs, sorted – directly comparable with plain tuple</==(same first-point-of- difference semantics ascompare_locant_sets, since every produced tuple here has the same shape: 2-tuples of ints, never a bare int or a fusion(int, str)locant).Shared by
bicyclo.py::get_bicyclo_numbering(bicyclic andpolycyclic.py::VonBaeyerAnalyzer._unsaturation_locant_key(tri-/polycyclic so the two engines apply one rule, not two.
- orthonym.rules.locants.compare_numbering(candidate_a, candidate_b)#
Fully-ordered deterministic numbering tie-break + layers).
The single shared lowest-locant comparator routed through by all three numbering engines (skeletal-replacement, fused-ring/PAH, benzene ring) so that the same structure always yields the same numbering regardless of input SMILES order. It NEVER falls through to input order: when every active tier ties, the two numberings are genuinely equivalent (the molecule has that symmetry) and either is correct & byte-identical.
Each candidate is a dict carrying only the tiers relevant to the call; a missing key means that tier is unconstrained (ties). Tiers, applied in the
order:
- ‘heteroatoms’ list of (locant, element) — positional set (b)
then element-seniority lowest-locant /
‘indicated_h’ indicated-hydrogen locant set (c)) ‘pcg’ principal-characteristic-group locant set (d)) ‘substituents’ detachable-prefix locant set (f)) ‘alpha’ sortable key giving the alphabetically-first (g))
prefix the lowest locant
Returns -1 if
candidate_ais preferred, +1 ifcandidate_b, 0 if every active tier ties.compare_locant_setsis reused unchanged as the positional primitive — this comparator is additive alongside it.
- orthonym.rules.locants.orient_chain(chain, mol, principal_group_atoms, double_bonds, triple_bonds, substituent_positions=None)#
Orient the principal chain to produce the lowest IUPAC locant set.
- Applies IUPAC 2013 criteria in strict order:
Lowest locants for principal characteristic group
Lowest locants for multiple bonds (double + triple, as a set)
Lowest locants for double bonds specifically
Lowest locants for substituents (detachable prefixes)
Lowest locant for alphabetically first substituent (g))
For pure hydrocarbons (no principal group), criterion (a) is skipped.
- Parameters:
chain (List[int]) – Ordered list of atom indices forming the candidate chain. This function evaluates forward vs reversed ordering.
mol – RDKit Mol object (used to detect bonds on the chain).
principal_group_atoms (Set[int]) – Set of atom indices belonging to the principal characteristic group (may be empty).
double_bonds (List[Tuple[int, int]]) – List of (atom_idx, atom_idx) tuples for C=C bonds.
triple_bonds (List[Tuple[int, int]]) – List of (atom_idx, atom_idx) tuples for C#C bonds.
substituent_positions (Dict[int, List] | None) – Optional dict mapping chain atom index to list of substituent groups. If None, criterion (d) is skipped.
- Returns:
The chain in the preferred orientation (may be the same list or reversed).
- Return type:
List[int]
- orthonym.rules.locants.get_functional_group_locants(chain, fg_atom_tuples, atom_to_locant, mol=None)#
Resolve functional group SMARTS match atoms to chain locants.
Given FG match tuples from SMARTS pattern matching, find the locant-defining atom for each match on the principal chain.
The locant-defining atom is the carbon on the chain that bears the functional group – e.g. the carbonyl carbon for ketones, the carbon bearing -OH for alcohols, the carboxyl carbon for acids.
When an RDKit mol is provided, the function selects the on-chain carbon with the most bonds to heteroatoms (O, N, S, etc.) within the match. This correctly handles SMARTS patterns where a neighbor carbon appears before the key carbon in the match tuple (e.g. the ketone pattern
[#6][CX3](=O)[#6]).Without a mol, falls back to picking the first on-chain atom.
- Parameters:
chain (List[int]) – Ordered principal chain atom indices.
fg_atom_tuples (List[Tuple[int, ...]]) – List of atom index tuples from SMARTS matching. Each tuple contains indices of atoms in one FG instance.
atom_to_locant (Dict[int, int]) – Mapping from atom index to locant (from build_atom_to_locant).
mol – Optional RDKit Mol object. When provided, used to score candidate carbon atoms by their heteroatom connectivity.
- Returns:
Sorted list of locants where the functional group attaches to the chain. Empty list if no FG atoms are on the chain.
- Return type:
List[int]
- orthonym.rules.locants.get_bond_locants(chain, bonds, atom_to_locant)#
Get locants for bonds (double or triple) within the chain.
For each bond (atom_a, atom_b), both atoms must be in the chain. The locant is the LOWER of the two atom locants (per IUPAC convention, a bond between positions i and i+1 is cited by locant i).
- Parameters:
chain (List[int]) – Ordered list of atom indices in the principal chain.
bonds (List[Tuple[int, int]]) – List of (atom_idx_a, atom_idx_b) tuples for bonds.
atom_to_locant (Dict[int, int]) – Mapping from atom index to locant (1-indexed).
- Returns:
Sorted list of locants for the bonds.
- Return type:
List[int]
Examples
>>> chain = [0, 1, 2, 3] >>> bonds = [(1, 2)] # bond between atoms 1 and 2 >>> atom_to_locant = {0: 1, 1: 2, 2: 3, 3: 4} >>> get_bond_locants(chain, bonds, atom_to_locant) [2]