orthonym.rules.heterocycles#

Note

Internal API. Names and behaviour may change between releases.

Heterocycle detection, classification, ring numbering, and naming according to IUPAC 2013.

Handles: - Heterocycle classification (ring size, aromaticity, heteroatom types) - Ring numbering starting at highest-priority heteroatom - Direction selection to minimize locants for other heteroatoms - First-point-of-difference rule for direction ties - Retained name lookup for common heterocycles - Hantzsch-Widman (HW) systematic name generation

IUPAC 2013 Rules for heterocycle numbering: - Position 1 goes to highest-priority heteroatom (O > S > Se > Te > N > P >…) - Numbering direction chosen to give lowest locants to other heteroatoms - First-point-of-difference comparison (not sum of locants)

Naming priority: 1. Check retained names FIRST (pyridine, furan, morpholine, etc.) 2. Fall back to HW systematic naming if no retained name

Reference: IUPAC 2013 Blue Book, Section (Heterocycles)

orthonym.rules.heterocycles.classify_heterocycle(mol, ring_atoms)#

Classify a heterocyclic ring by its structural features.

Returns detailed information about the ring for naming purposes: - Ring size (number of atoms) - Heteroatom identity and positions - Aromaticity - Saturation status - Dominant heteroatom (highest priority)

Parameters:
  • mol – RDKit Mol object

  • ring_atoms – Iterable of atom indices defining the ring

Returns:

Dict with –

  • ring_size: int (3-10 typically)

  • heteroatoms: List[Tuple[int, str]] - (atom_idx, element_symbol)

  • is_aromatic: bool

  • is_saturated: bool (no double bonds within ring)

  • num_heteroatoms: int

  • dominant_heteroatom: str (highest priority element in ring)

Return type:

Dict

Examples

>>> mol = Chem.MolFromSmiles('c1ccncc1') # pyridine
>>> info = classify_heterocycle(mol, mol.GetRingInfo.AtomRings[0])
>>> info['ring_size']
6
>>> info['is_aromatic']
True
>>> info['dominant_heteroatom']
'N'
orthonym.rules.heterocycles.number_heterocycle_ring(mol, ring_atoms, heteroatoms=None)#

Reorder ring atoms for IUPAC numbering of heterocycles.

Position 1 is assigned to the highest-priority heteroatom. Numbering direction is chosen to give lowest locants to other heteroatoms, using first-point-of-difference comparison.

Parameters:
  • mol – RDKit Mol object

  • ring_atoms – Iterable of atom indices defining the ring

  • heteroatoms (List[Tuple[int, str]] | None) – Optional list of (atom_idx, element) tuples. If None, will be computed from ring_atoms.

Returns:

List of ring atom indices reordered for IUPAC numbering. Position 1 (index 0) is the highest-priority heteroatom.

Return type:

List[int]

Examples

>>> mol = Chem.MolFromSmiles('c1ccncc1') # pyridine
>>> ring = mol.GetRingInfo.AtomRings[0]
>>> oriented = number_heterocycle_ring(mol, ring)
>>> # N atom should be at position 1 (index 0)
>>> mol.GetAtomWithIdx(oriented[0]).GetSymbol
'N'
orthonym.rules.heterocycles.orient_heterocycle(mol, ring_atoms)#

Convenience function combining classification and numbering.

Parameters:
  • mol – RDKit Mol object

  • ring_atoms – Iterable of atom indices defining the ring

Returns:

Tuple of –

  • oriented_ring: List of atom indices in IUPAC numbering order

  • atom_to_locant: Dict mapping atom_idx to locant (1-indexed)

Return type:

Tuple[List[int], Dict[int, int]]

Examples

>>> mol = Chem.MolFromSmiles('c1ccncc1') # pyridine
>>> ring = mol.GetRingInfo.AtomRings[0]
>>> oriented, mapping = orient_heterocycle(mol, ring)
>>> # N is at locant 1
>>> n_idx = [i for i in oriented if mol.GetAtomWithIdx(i).GetSymbol == 'N'][0]
>>> mapping[n_idx]
1
orthonym.rules.heterocycles.orient_heterocycle_with_substituents(mol, ring_atoms, substituent_positions=None, principal_group_atoms=None)#

Orient heterocycle considering heteroatoms, the principal characteristic group, AND other substituent positions.

IUPAC Rule low-locant order, applied to a ring whose senior heteroatom is fixed at position 1 per: after fixing the heteroatom at position 1, choose the numbering direction that gives lowest locants to, IN THIS ORDER: 1. Other heteroatoms (if any) 2. The principal characteristic group (suffix) [(c)] 3. Other substituents (if the above tie) [(g)]

principal_group_atoms is the set of RING atom indices that bear the principal characteristic group (e.g. the ring carbon double-bonded to =O for a ketone, or the ring carbon bearing -OH for an -ol). When None/empty the behaviour is exactly the prior heteroatom-then-substituent order.

Parameters:
  • mol – RDKit Mol object

  • ring_atoms – Iterable of atom indices defining the ring

  • substituent_positions (Set[int] | None) – Set of ring atom indices that have substituents. If None, will be detected from the molecule.

Returns:

Tuple of –

  • oriented_ring: List of atom indices in IUPAC numbering order

  • atom_to_locant: Dict mapping atom_idx to locant (1-indexed)

Return type:

Tuple[List[int], Dict[int, int]]

Examples

>>> mol = Chem.MolFromSmiles('Cc1ccccn1') # 2-methylpyridine
>>> ring = mol.GetRingInfo.AtomRings[0]
>>> # Find which ring atom has the methyl substituent
>>> sub_positions = {idx for idx in ring if _has_substituent(mol, idx, ring)}
>>> oriented, mapping = orient_heterocycle_with_substituents(mol, ring, sub_positions)
>>> # Should give methyl the lowest possible locant (2, not 6)
orthonym.rules.heterocycles.get_heteroatom_locants(oriented_ring, mol)#

Get locants and element symbols for all heteroatoms in an oriented ring.

Parameters:
  • oriented_ring (List[int]) – Ring atoms already in IUPAC numbering order (position 1 at index 0)

  • mol – RDKit Mol object

Returns:

List of (locant, element_symbol) tuples for each heteroatom, sorted by locant.

Return type:

List[Tuple[int, str]]

Examples

>>> mol = Chem.MolFromSmiles('c1ncnc1') # pyrimidine-like
>>> ring = mol.GetRingInfo.AtomRings[0]
>>> oriented, _ = orient_heterocycle(mol, ring)
>>> locants = get_heteroatom_locants(oriented, mol)
>>> # Returns something like [(1, 'N'), (3, 'N')]
orthonym.rules.heterocycles.get_ring_canonical_smiles(mol, ring_atoms)#

Extract a ring as canonical SMILES for retained name lookup.

Creates a new molecule containing only the ring atoms and their bonds, then returns the canonical SMILES representation.

MolFragmentToSmiles writes each atom with the hydrogen count it has in mol, so a ring whose indicated hydrogen has been replaced by a substituent comes back as a string that describes no molecule at all (Cn1cccc1 -> c1ccnc1, unkekulisable). Such a string can never equal a table key, and the retained name is lost purely to the extraction. When – and only when – that happens, the displaced indicated hydrogen is restored and the fragment rewritten (_ring_smiles_with_indicated_h_restored).

The repair is therefore strictly additive: a ring whose fragment already parses is returned exactly as before, so no key that resolves today can change. The only outcomes that move are keys that match nothing.

Parameters:
  • mol – RDKit Mol object

  • ring_atoms – Iterable of atom indices defining the ring

Returns:

Canonical SMILES string for the ring

Return type:

str

Examples

>>> mol = Chem.MolFromSmiles('c1ccncc1') # pyridine
>>> ring = mol.GetRingInfo.AtomRings[0]
>>> get_ring_canonical_smiles(mol, ring)
'c1ccncc1'
>>> mol = Chem.MolFromSmiles('Cn1cccc1') # 1-methyl-1H-pyrrole
>>> ring = mol.GetRingInfo.AtomRings[0]
>>> get_ring_canonical_smiles(mol, ring)
'c1cc[nH]c1'
orthonym.rules.heterocycles.build_hw_name(heteroatoms, ring_size, is_saturated, is_aromatic, lambda_by_locant=None, cite_locants=True)#

Build Hantzsch-Widman systematic name for a heterocycle.

cite_locants=False leaves out the heteroatom locant set, for a caller that has established a licence for the whole name (a lambda locant is always cited,.

Returns None – fail closed – when any heteroatom has no Table 2.4 prefix. "" still means “no heteroatoms supplied” and stays distinct.

Assembles a systematic HW name from heteroatom prefixes and ring stem: 1. Group heteroatoms by element 2. Order by IUPAC priority (O > S > N >…) 3. Add multipliers and locants for multiple same heteroatoms 4. Get stem based on ring size and saturation 5. Apply ‘a’ elision (drop terminal ‘a’ before vowel stem)

Parameters:
  • heteroatoms (List[Tuple[int, str]]) – List of (locant, element) tuples for heteroatoms in numbered order

  • ring_size (int) – Ring size (3-10)

  • is_saturated (bool) – True if fully saturated

  • is_aromatic (bool) – True if aromatic (overrides is_saturated for stem selection)

Returns:

HW systematic name (e.g., ‘oxolane’, ‘1,3-dioxolane’, ‘azine’)

Return type:

str | None

Examples

>>> build_hw_name([(1, 'O')], 5, True, False)
'oxolane'
>>> build_hw_name([(1, 'O'), (3, 'O')], 5, True, False)
'1,3-dioxolane'
>>> build_hw_name([(1, 'N')], 6, False, True)
'azine'
orthonym.rules.heterocycles.name_partially_saturated_monocyclic_heterocycle(mol, ring_atoms, principal_group_atoms=None)#

Name a partially-saturated monocyclic mancude heterocycle (IUPAC.

Example: C1C=CC=CN1 -> 1,2-dihydropyridine (the systematic HW path would otherwise drop the hydrogenation and emit the mancude parent pyridine). Works by (1) building the mancude aromatic parent and naming it, (2) finding the ring atoms that carry a ring double bond in the mancude parent but have LOST it in the molecule (the hydro positions — detected by bond topology, NOT hybridization, so a conjugated enamine N still counts), (3) numbering the ring with the heteroatom set lowest then the hydro set lowest, and (4) emitting <locants>-<prefix>hydro<parent>.

Scope (fail closed otherwise — never a wrong name, -safe):
  • the mancude parent must aromatize and carry NO leading indicated hydrogen (parents like 1H-pyrrole / 2H-pyran whose hydro form interleaves added-indicated-H are deferred to a follow-on);

  • the molecule must retain >= 1 ring unsaturation (a fully saturated ring is named by its retained / HW saturated stem, not as a hydro prefix);

  • the hydro count must be a standard di/tetra/hexa… value.

orthonym.rules.heterocycles.name_heterocycle(mol, ring_atoms, principal_group_atoms=None)#

Generate IUPAC name for a heterocyclic ring.

Returns None when a ring heteroatom has no replacement prefix in the governing table, so no name can express it (see build_hw_name).

Naming priority: 1. Check retained names FIRST (pyridine, furan, morpholine, etc.) 2. For rings 3-10: Hantzsch-Widman systematic naming 3. For rings > 10: replacement (“a”) nomenclature on cycloalkane parent

Parameters:
  • mol – RDKit Mol object

  • ring_atoms – Iterable of atom indices defining the ring

Returns:

IUPAC name for the heterocycle

Return type:

str | None

Examples

>>> mol = Chem.MolFromSmiles('c1ccncc1') # pyridine
>>> ring = mol.GetRingInfo.AtomRings[0]
>>> name_heterocycle(mol, ring)
'pyridine'
>>> mol2 = Chem.MolFromSmiles('C1CO1') # oxirane
>>> ring2 = mol2.GetRingInfo.AtomRings[0]
>>> name_heterocycle(mol2, ring2)
'oxirane'
orthonym.rules.heterocycles.ring_principal_suffix_atoms(mol, ring_atoms, principal_group, functional_groups=None)#

RING atom indices that carry the principal characteristic group as a SUFFIX on this heterocycle parent, so numbering gives them the lowest locants.

IUPAC NUMBERING (the Blue Book): “low locants are assigned to them in the following decreasing order of seniority” —… (c) “principal characteristic groups and free valences (suffixes)”;… (f) “detachable alphabetized prefixes”. Criterion (c) — the -dione / -sulfonic acid SUFFIX — outranks (f), the methyl prefix. (For a heterocycle the ring heteroatoms are numbered first, change note (1) at the Blue Book; the heteroatom set ties in both directions in these cases, so the suffix decides.) orient_heterocycle_with_substituents receives this set as principal_group_atoms.

The generic prefix-class anchor in namer.py keys on get_prefix(principal_group) and only anchors a CARBON-centred appended suffix, so it MISSES two families whose suffix decision lives in get_heterocycle_substituents below:

  1. Pseudoketone ring carbons named -one/-dione/-thione — a ring carbon bearing an exocyclic =O/=S/=Se/=Te whose molecule-level principal group is a lactam-declined amide, a cyclic imide, a lactone-declined ester, or a thioamide. get_prefix of those principal groups is None (or carbamoyl), never oxo, so the ring carbonyl carbons were never anchored and the (f) substituent set put a SUBSTITUTED ring atom at locant 1 (CN1C(=O)CNC1=O -> 1-methyl...-2,5-dione where (c) requires 3-methyl...-2,4-dione). Mirrors the ring_ketone_suffix / ring_imide_suffix / ring_lactone_suffix / ring_thioamide_suffix gates in get_heterocycle_substituents exactly, but needs no orientation (those gates are locant-independent).

  2. The ring atom (typically the ring N) bearing an exocyclic SULFUR-oxoacid / sulfonamide appended suffix (-sulfonic acid / -sulfinic acid / -sulfonamide). Its central atom is sulfur, not carbon, so namer.py’s carbon-only anchor skipped it and the substituted ring N took locant 1 (OS(=O)(=O)N1CCN(C)CC1 -> 1-methylpiperazine-4-sulfonic acid where (c) requires 4-methylpiperazine-1-sulfonic acid). The ring atom bonded to the suffix sulfur is the anchor — the sulfur analogue of the ring-N carboxylic-acid anchor the carbon path already feeds.

orthonym.rules.heterocycles.get_heterocycle_substituents(mol, ring_atoms, oriented_ring, atom_to_locant, principal_group=None)#

Find substituents attached to a heterocyclic ring.

For each ring atom that has neighbors not in the ring, identifies the substituent and tracks whether it’s attached to a nitrogen (N-substitution). Also classifies substituents as ring or alkyl to correctly name ring substituents (piperidinyl, phenyl) instead of counting carbons (pentyl, hexyl).

Parameters:
  • mol – RDKit Mol object

  • ring_atoms – Iterable of atom indices defining the ring

  • oriented_ring (List[int]) – Ring atoms in IUPAC numbering order

  • atom_to_locant (Dict[int, int]) – Dict mapping atom_idx to locant (1-indexed)

Returns:

Dict mapping locant to list of substituent info dicts –

{
locant: [
{

‘atoms’: List[int], # Atom indices in substituent ‘is_on_nitrogen’: bool, # True if attached to N ‘carbon_count’: int, # Number of C atoms (for alkyl naming) ‘connecting_atom’: int, # Ring atom the sub attaches to ‘is_ring’: bool, # True if substituent is a ring ‘ring_name’: str, # Ring substituent name (if is_ring)

}

]

}

Return type:

Dict[int, List[Dict]]

Examples

>>> mol = Chem.MolFromSmiles('CN1CCCC1') # N-methylpyrrolidine
>>> ring = mol.GetRingInfo.AtomRings[0]
>>> oriented, atom_to_loc = orient_heterocycle(mol, ring)
>>> subs = get_heterocycle_substituents(mol, ring, oriented, atom_to_loc)
>>> # N-methyl should be at locant 1 (N position) with is_on_nitrogen=True
orthonym.rules.heterocycles.name_substituted_heterocycle(mol, ring_atoms, parent_name, substituents, atom_to_locant, principal_group=None)#

Assemble complete name for a substituted heterocycle.

Returns None (the heterocycle candidate declines) when an exocyclic substituent cannot be named correctly and completely — naming it partially would silently drop atoms (sugar-ring-oxygen-drop / structure loss). Callers treat a None return as a fail-closed decline.

N-substituted groups use N-locant format (N-methyl, N,N-dimethyl). C-substituted groups use numeric locants (2-methyl, 3-ethyl). Ring substituents use proper ring names (piperidinyl, phenyl) not carbon counts. All prefixes are sorted alphabetically (ignoring N-, numbers, multipliers). Stereochemistry descriptors (R/S) are added as a prefix if present.

Parameters:
  • mol – RDKit Mol object

  • ring_atoms – Iterable of atom indices defining the ring

  • parent_name (str) – Base heterocycle name (e.g., ‘pyrrolidine’, ‘pyridine’)

  • substituents (Dict[int, List[Dict]]) – Dict from get_heterocycle_substituents

  • atom_to_locant (Dict[int, int]) – Dict mapping atom_idx to locant (1-indexed)

Returns:

Complete IUPAC name (e.g., ‘N-methylpyrrolidine’, ‘3-methylpyridine’, ‘(2S)-N-methyl-2-propylpiperidine’)

Return type:

str | None

Examples

>>> mol = Chem.MolFromSmiles('CN1CCCC1') # N-methylpyrrolidine
>>> #... get ring, orient, get substituents...
>>> name_substituted_heterocycle(mol, ring, 'pyrrolidine', subs, atom_to_loc)
'N-methylpyrrolidine'
orthonym.rules.heterocycles.get_saturation_prefix(mol, ring_atoms, is_aromatic_parent)#

Get saturation prefix for partially saturated heterocycles (HETERO-08).

For saturated versions of aromatic parents, returns prefixes like: - dihydro- (2 additional H atoms) - tetrahydro- (4 additional H atoms)

Note: This is a placeholder - full implementation requires comparing to the aromatic parent structure.

Parameters:
  • mol – RDKit Mol object

  • ring_atoms – Iterable of atom indices defining the ring

  • is_aromatic_parent (bool) – True if the parent is aromatic (e.g., furan vs THF)

Returns:

Saturation prefix string (e.g., ‘tetrahydro-’) or empty string

Return type:

str

Examples

>>> # Tetrahydrofuran is "oxolane" (retained name), but could also be
>>> # named "tetrahydrofuran" relating it to furan
>>> # This function helps determine the saturation prefix when needed