orthonym.assembly.substituent_enumerator#

Note

Internal API. Names and behaviour may change between releases.

Unified substituent enumeration module with ReplaceCore-based extraction.

Provides a single enumeration path for all substituents on both ring and chain parent structures. Replaces the previously fragmented three-path system that caused silent drops, double-counting, and wrong locants.

Architecture:
  • Ring parents: ReplaceCore(mol, core_from_ring_atoms) -> fragment mols

  • Chain parents: Branch-point enumeration from features.substituents dict

  • All fragments: classify -> name via existing naming infrastructure

Public API:
  • discover_substituents(mol, parent_atoms, parent_type,…) [a phase]

  • extract_ring_substituents(mol, ring_atoms, oriented_ring)

  • extract_chain_substituents(mol, principal_chain, substituents_dict)

  • classify_and_name_fragment(mol, frag_info, parent_atoms, features=None)

  • collect_substituent_atom_set(substituent_infos)

References

IUPAC 2013 (detachable prefixes) IUPAC 2013 (parent selection determines what’s a substituent)

class orthonym.assembly.substituent_enumerator.SubstituentInfo(frag_mol, locant, attach_mol_idx, frag_atoms)#

Bases: tuple

Represents a single substituent on a parent structure.

Fields:
frag_mol: RDKit Mol of the isolated fragment (with dummy atom at attachment),

or None for chain-parent substituents where fragment is inline.

locant: IUPAC locant (1-indexed integer) on the parent. attach_mol_idx: Original mol atom index of the attachment point on parent. frag_atoms: Set of original mol atom indices belonging to this substituent.

attach_mol_idx#

Alias for field number 2

frag_atoms#

Alias for field number 3

frag_mol#

Alias for field number 0

locant#

Alias for field number 1

orthonym.assembly.substituent_enumerator.discover_substituents(mol, parent_atoms, parent_type='auto', oriented_ring=None, principal_chain=None, atom_to_locant=None, ring_atom_to_locant=None, general_fallback=False)#

Discover ALL substituents on a parent structure.

Universal entry point that replaces six parallel substituent discovery systems. Every non-parent, non-hydrogen atom in mol is assigned to exactly one SubstituentInfo. No silent drops, no size limits.

Parameters:
  • mol – RDKit Mol object.

  • parent_atoms – Set of atom indices defining the parent structure.

  • parent_type – "ring", "chain", or "auto" (auto-detects).

  • oriented_ring – Ring atom indices in IUPAC order (required for ring parents).

  • principal_chain – Chain atom indices in order (required for chain parents).

  • atom_to_locant – Optional mapping of atom idx -> IUPAC locant. Consumed by the CHAIN path only.

  • ring_atom_to_locant – Optional mapping of atom idx -> IUPAC locant for a RING parent, inherited from whichever producer spelled the parent name. Overrides the oriented_ring position arithmetic per atom ('4a'/'8a' and any fused-parent numbering cannot be expressed as a position). None -> unchanged behaviour.

  • general_fallback –

    Composer1 Task 5 gating flag. Passed True ONLY from the general-engine (complete/best-effort) call sites. Controls the failure mode when a non-parent heavy atom cannot be assigned to any substituent fragment — e.g. a substituent hanging off a SUFFIX/FG heteroatom (the N-aryl ring of an amide anilide), which the parent- chain/ring walk cannot reach:

    • False (PIN default): _verify_completeness HARD-asserts on the unassigned atoms exactly as before (byte-identical). The narrow PIN handlers own these molecules; a partition gap there is a genuine invariant violation the PIN callers already catch.

    • True (general fallback): FAIL CLOSED to a clean sentinel — return None instead of raising, and NEVER silently drop the unassigned atoms (dropping them would ship a name of the wrong constitution). The general caller converts the None to a clean _refuse -> the engine abstains, no exception.

Returns:

List[SubstituentInfo] with one entry per substituent fragment, or None when general_fallback=True and the partition is incomplete (unassigned / overlapping atoms) — signalling the caller to fail closed.

class orthonym.assembly.substituent_enumerator.FreeValencePrefix(free_valence, prefix, basis, in_class=True)#

Bases: object

The verdict for one substituent fragment’s attachment bond.

free_valence is always the TRUE bond order at the attachment (1/2/3), or None when the shape is outside the single-attachment-atom case a free-valence prefix describes at all (a bridge, an aromatic linkage). in_class is False when the attachment atom is not carbon, where the

carbon morphology does not apply and the caller’s own heteroatom

machinery owns the naming.

Callers branch on the two PROPERTIES, never on the fields: reading the fields directly is how a caller re-invents the policy this class holds – and getting in_class wrong turns every ring ketone into an abstention.

free_valence: int | None#
prefix: str | None#
basis: str#
in_class: bool = True#
property defers: bool#

True when the caller’s own existing naming path is -correct.

property must_fail_closed: bool#

True when needs a multivalent prefix and none could be built.

orthonym.assembly.substituent_enumerator.carbon_free_valence_prefix(mol, frag_atoms, attach_idx)#

verdict for a CARBON-attached substituent fragment.

The shared entry point for every substituent detector – monocyclic, heterocyclic, bicyclo, von Baeyer polycyclic, spiro, fused. Use it wherever a fragment is about to be named from its carbon count:

verdict = carbon_free_valence_prefix(mol, sub_atoms, first_atom)
if verdict.prefix:
    name = verdict.prefix
elif verdict.must_fail_closed:
    ...abstain... # never get_alkyl_name(carbon_count)
else:
    name = get_alkyl_name(carbon_count) # unchanged legacy path
Parameters:
  • mol – RDKit Mol of the FULL molecule.

  • frag_atoms – Atom indices of the substituent fragment.

  • attach_idx – The FRAGMENT atom bonded to the parent (not the parent atom – the linkage bond runs from this atom OUT of the fragment).

Scope – deliberately CARBON only. The multivalent heteroatom prefixes (oxo, sulfanylidene, imino) are a separate class that the detectors’ own heteroatom machinery already names correctly, and their tokens spell no morpheme at all, so the text check below cannot confirm them. Including them would turn every ring ketone into a fail-closed abstention. A non-carbon attachment therefore DEFERS, leaving the existing chalcogen/nitrogen paths untouched.

The prefix is accepted only when its own TEXT spells the number of free valences the BOND has. The two facts come from independent places – the count from the molecular graph, the reading from the shared morpheme oracle that the proof spine’s P7 also uses – so their agreement is evidence, and a future constructor change that quietly emitted a -yl token would be caught here rather than shipped.

Deliberately calls the CONSTRUCTOR, not the full naming cascade. Two reasons: no cascade tier mints a carbon -ylidene token (they all build single-valence prefixes, which the check would reject anyway), and this function is called FROM inside the cascade’s own chokepoints, so routing it back through them would recurse without terminating.

References: IUPAC 2013,,.

orthonym.assembly.substituent_enumerator.name_substituent_for_ordering(mol, frag_atoms, attach_idx)#

name_substituent() for a SORT KEY only – it leaves no trace.

Memoised for the top-level naming call (assembly.memo; verify mode recomputes and compares): a ring’s (g) tie is re-asked for the same prefix every time the ring is oriented, and each ask re-named that prefix from scratch – 3,458 sort-key namings (92 of 141 s) for a ChEBI siderophore whose piperazinedione carries two long hydroxamate chains (a NATIVE_HANG row; the measured code named it in 17 s). The key is the whole atom-ordered graph plus the fragment and its attachment atom, i.e. everything the call reads.

A numbering tie-break (g): the prefix cited first takes the lower locant) needs each prefix’s name BEFORE the real assembly names it. A plain name_substituent call there is not free of side effects: it writes the scope memo, the fragment memo cache (which is context-dependent), the per-molecule polyfunctional memo and the molecule’s SMILES output-order props, and it spends the work budgets – measured: one such call changed a later, unrelated prefix name of the same molecule. This wrapper runs the call against copies of all of them and restores them, so the real naming is exactly what it would have been without the speculative call. Returns None for the unnameable sentinel.

orthonym.assembly.substituent_enumerator.name_substituent(mol, frag_atoms, attach_idx, allow_mancude=False)#

Name a substituent fragment with the free-valence morphology requires.

Thin gate over:func:_name_substituent_cascade, which holds the five-tier naming logic. The split exists because the cascade’s tiers all build their token from the FRAGMENT alone – they never see the attachment bond – so none of them can know whether the free valence is single (-yl), double (-ylidene) or triple (-ylidyne). This function is the one place that holds the fragment and the attachment atom together, so the morphology is decided here, once, for every tier.

Behaviour by free valence:

  • one (or undecidable – a bridge, an aromatic linkage): the cascade runs and its answer is returned untouched. This is the overwhelmingly common case and is byte-identical to the pre- behaviour.

  • two or three: the carbon ylidene/ylidyne constructor gets first refusal; if it declines, the cascade runs and its answer is accepted ONLY if the token does not spell a single free valence. A -yl token on a doubly-bonded free valence names a DIFFERENT molecule, so it is refused rather than shipped – None under allow_mancude (the Tier-4.5 de-masking convention) and the 'substituent' unnameable sentinel otherwise, both of which fail closed downstream.

The correct multivalent heteroatom prefixes (oxo, sulfanylidene, imino,…) spell no -yl morpheme, so they pass the guard untouched.

References

IUPAC 2013 (free-valence morphology), (lowest locant for the free valence).

orthonym.assembly.substituent_enumerator.name_ylidene_substituent(mol, frag_atoms, attach_idx)#

Name a fragment whose bond to its parent is DOUBLE, or None.

The shared entry point for the namers that cite a doubly-bonded fragment – hydrazone, semicarbazone, azine, the cumulative ium/ide chain, the (3) ketene branch. Each of them used to spell the morphology itself:

yl = name_substituent(mol, frag, c)
if not yl.endswith("yl"):
    return None
... f"{yl}idene"...

which decided the free valence a second time, in the consumer, by rewriting a token. It only worked while the pipeline was returning the WRONG single-valence token for a double bond; correcting the pipeline made every such consumer reject its own correct input.

Here the producer owns the morphology and the consumer VERIFIES it: the token comes back from name_substituent already carrying the ending its attachment bond earned, and is returned only if its own text confirms the two free valences. Nothing is appended, so there is no second place for the two to disagree.

orthonym.assembly.substituent_enumerator.name_ylidyne_substituent(mol, frag_atoms, attach_idx)#

Name a fragment whose bond to its parent is TRIPLE, or None.

The -ylidyne sibling of:func:name_ylidene_substituent: for a fragment attached by a triple bond (three free valences), e.g. the CH3-C# of a nitrile imide CC#[N+][N-]C -> ethylidyne. As with the ylidene entry point, the token comes back from:func:name_substituent already carrying the morphology its attachment bond earned, and is returned only if its own text confirms the three free valences – nothing is appended, so producer and consumer cannot drift about the free valence.

orthonym.assembly.substituent_enumerator.composed_prefix_organyl_name(mol, frag_atoms, attach_idx)#

The PIN substituent-prefix token for an organyl fragment, or None.

THE organyl-naming step for every composed heteroatom prefix, and the root-cause replacement for get_alkyl_name(carbon_count) at those sites. The fragment is routed through the shared audited cascade (name_substituent()), which PERCEIVES the branching a count erases.

Strict acceptance filter, following the Phase-3 substituent_purity.organyl_prefix_name pattern: a single prefix TOKEN or nothing. Every refusal sentinel is recognised by the ONE shared errors.is_refusal_sentinel() predicate – a sentinel welded into a prefix slot becomes part of a name that reads as success – and a multi-word result is refused too, because a composed prefix must never be assembled around a whole compound name.

Pure: no mol mutation.

orthonym.assembly.substituent_enumerator.cite_organyl_in_composed_prefix(token)#

Cite token inside a composed prefix with its enclosing marks.

A SIMPLE prefix is cited BARE: ‘butylamino’ and ‘methylsulfanyl’ (BB 18089 “-NH-CH3 methylamino (preferred prefix)”; BB 27651 “CH3-S- methylsulfanyl (preferred prefix)”). A retained ITALICISED prefix is simple too – ‘tert-butyl’ is a preferred prefix cited bare in ‘tert-butyldi(methyl)phosphane (PIN)’, BB 16286).

A COMPOUND prefix takes enclosing marks with the composing suffix OUTSIDE them: ‘(propan-2-yl)oxy’ and ‘(butan-2-yl)oxy’ are the preferred prefixes , BB 27683/27687) and ‘(chloromethyl)amino’ likewise (BB 18112). The marks go around the ORGANYL, never around the whole composed prefix.

The italicised-prefix carve-out is decided by the ONE shared naming_utils.italicized_prefix_is_bare(), never re-spelled here.

orthonym.assembly.substituent_enumerator.composed_prefix_multiplier(token, count)#

The multiplicative prefix for count copies of token, or None.

(a) / (a). A SIMPLE substituent prefix – including one that merely carries a locant – takes ‘di’/’tri’: BB 25719 ‘1,4-di(propan-2-yl)cyclohexane (PIN)’ and BB 28170 ‘2-[di(butan-2-yl)amino]butan-2-ol (PIN)’. A COMPOUND (substituted) prefix takes ‘bis’/’tris’: BB 41118 ‘bis(2-methylpropyl)’, BB 40703 ‘bis(chloromethyl)aminoxyl (PIN)’ and BB 7104 ‘bis(2-chloropropan-2-yl) (preferred prefix)’.

None when the multiplicity has no tabulated prefix – the caller must then fail closed rather than invent one.

orthonym.assembly.substituent_enumerator.composed_alkoxy_prefix(token)#

Turn an organyl token into its alkoxy prefix, or None.

The Blue Book tabulates this morphology verbatim (BB 27667-27691):

  • CH3-[CH2]3-O- butoxy – the retained CONTRACTED prefixes (methoxy, ethoxy, propoxy, butoxy) are simple and cited bare;

  • (CH3)3C-O- *tert*-butoxy “(preferred prefix) (no substitution)”, explicitly NOT ‘tert-butyloxy’ (BB 55662);

  • (CH3)2CH-O- (propan-2-yl)oxy and CH3-CH2-CH(CH3)-O- (butan-2-yl)oxy – a locant-bearing free valence keeps the alkyl name whole inside marks, suffix outside;

  • (CH3)2CH-CH2-O- 2-methylpropoxy “(preferred prefix) (not isobutoxy)” – a free valence at position 1 CONTRACTS even when the chain is branched, and is cited bare.

None when the token is not an alkyl the table covers, so the caller fails closed instead of guessing a morphology.

orthonym.assembly.substituent_enumerator.alkoxy_prefix_from_substituent(token)#

A ‘…yl’ substituent NAME -> its R-oxy prefix.

Thin wrapper over:func:composed_alkoxy_prefix that ALSO applies the decorated-phenyl -> phenoxy contraction the primitive deliberately declines. composed_alkoxy_prefix stays conservative on any ‘…phenyl’ token (the biphenyl ‘4-phenylphenyl’ shape, BB 24607 -> ‘([1,1’-biphenyl]-4-yl)oxy’) and leaves the contraction to the caller. A DECORATED benzene is the retained, fully-substitutable ‘phenoxy’ / BB 17796 “phenoxy… full substitution”): ‘4-methylphenyl’ -> ‘4-methylphenoxy’. A locant-bearing biphenyl reaches us as ‘[1,1’-biphenyl]-4-yl’ (a ‘-N-yl’ token, routed by composed_alkoxy_prefix’s locant branch), never as ‘…phenyl’, so contracting a bare ‘…phenyl’ here is safe. Use this at every alkoxy emitter that hands in a general substituent name (F-spell-oxy).

orthonym.assembly.substituent_enumerator.composed_chalcogen_group_prefix(mol, frag_atoms, chalcogen_idx, boundary)#

The PIN prefix token for a MONO-chalcogen-rooted group -X-R, or None.

chalcogen_idx is a divalent O/S/Se/Te whose single continuation inside frag_atoms (not crossing boundary) is an ORGANYL group. The Blue Book spells both morphologies verbatim:

  • -O-CH3 -> methoxy, BB 27671)

  • -O-C(CH3)3 -> tert-butoxy (BB 27679, “not tert-butyloxy”)

  • -O-CH(CH3)2 -> (propan-2-yl)oxy (BB 27683)

  • -S-CH3 -> methylsulfanyl (BB 27651, BB 25021)

  • -Se-CH3 -> methylselanyl

THIS IS THE PRIMITIVE THE GENERIC CASCADE CANNOT SUPPLY. Handed a chalcogen attachment,:func:name_substituent RE-ROOTS the fragment at a carbon and names the chalcogen as a hydroxy/sulfanyl SUBSTITUENT on that carbon, so tert-Bu-O- came back as 2-hydroxy-2-methylpropyl – a different CONSTITUTION, not merely a different spelling. Any caller holding a chalcogen attachment must come here rather than guess with the cascade.

Scope is deliberately the MONO chalcogen. A second chalcogen beyond chalcogen_idx is the peroxy / disulfanyl class, which the cascade already names correctly and whole (‘methylperoxy’, ‘tert-butyldisulfanyl’); this returns None there so the caller keeps that working path.

Fails closed (None) when the atom is not a divalent chalcogen, has other than exactly one in-fragment continuation, that continuation is not carbon, or the organyl half cannot be named as a single prefix token.

orthonym.assembly.substituent_enumerator.chalcogen_rooted_acyl_amino_core(mol, carbonyl_c, n_idx, frag_atoms)#

The prefix core for a CHALCOGEN-rooted acyl on nitrogen, R-X-CO-NH-.

Returns '(tert-butoxycarbonyl)amino' and friends WITHOUT the outer enclosing marks (each caller applies its own, since the two call sites wrap at different depths), None when the acyl is not chalcogen-rooted so the caller keeps its carbon path, or:data:ACYL_CHALCOGEN_UNNAMEABLE when it IS this class but cannot be spelled – the caller must then fail closed.

THE CLASS THE COUNTS CANNOT REACH. R-O-CO- is an ester of carbamic acid (Boc, Cbz, Fmoc, methoxycarbonyl,…), not an acyl of a carboxylic acid, so no carbon count describes it. Both count-based callers proved that: one stops at the heteroatom and counts only the carbonyl carbon, returning 1 for a true formyl H-CO-NH- and for (CH3)3C-O-CO-NH- alike; the other WALKS ACROSS the heteroatom and counts the organyl beyond it as if it were acyl carbon. The first spelled Boc ‘methanoylamino’, the second spelled CH3-S-CO-NH- ‘ethanoylamino’.

The Blue Book builds this prefix by CONCATENATION onto the chalcogen-group prefix, verbatim:

-CO-O-CH2-C6H5 (benzyloxy)carbonyl (preferred prefix) [BB 18116]
CH3-CO-S-CO- (acetylsulfanyl)carbonyl (preferred prefix) [BB 18128]

and (BB 31698) names -CO-OR' ‘alkoxycarbonyl’. Whether the chalcogen prefix is cited BARE or in marks is the /

simple-vs-compound distinction – the retained contractions are

SIMPLE (BB 27667 “considered as simple prefixes”), so tert-butoxy concatenates bare, BB 54417 N2-(tert-butoxycarbonyl)-L-lysine, while a concatenated benzyloxy is COMPOUND (BB 27633) and takes its marks, BB 54422 N5-acetyl-N2-[(benzyloxy)carbonyl]-L-glutamine (both. That decision is NOT re-spelled here: it is the ONE shared enclose_if_compound, so these marks cannot drift from the rest of the system.

amino is the morpheme (BB 26314); with the caller’s marks the result is the Blue Book’s own [(acyl)amino]acetic acid shape, BB 33213 [(methanesulfinothioyl)amino]acetic acid (PIN).

Pure: no mol mutation.

orthonym.assembly.substituent_enumerator.is_dichalcogen_bridge_attach(mol, attach_idx, frag_atoms_set, require_different=False)#

True when attach_idx is a divalent chalcogen bonded, inside the fragment, to a second divalent chalcogen – the / bridge shapes -OO-, -SS-, -OS-, -SO-, -OSe-…

require_different=True narrows it to the MIXED bridge only (the two chalcogens are different elements), which is the shape the generic substituent tiers must never be allowed to mangle.

One walker for both questions on purpose: the two predicates differ by a single element comparison, and spelling them separately is how this class regrew a site at a time before.

orthonym.assembly.substituent_enumerator.extract_ring_substituents(mol, ring_atoms, oriented_ring, atom_to_locant=None)#

Extract all substituent fragments from a ring parent using ReplaceCore.

Builds a core mol from ring_atoms, calls ReplaceCore to extract all non-ring fragments as separate mol objects with isotope-labeled dummy atoms indicating attachment points.

Parameters:
  • mol – RDKit Mol object.

  • ring_atoms – Tuple or list of ring atom indices (from principal_ring).

  • oriented_ring – List of ring atom indices in IUPAC numbering order.

  • atom_to_locant –

    Optional {atom idx -> IUPAC locant} INHERITED from the producer that spelled the parent name. Where it covers an attachment atom it decides that atom’s locant; elsewhere the oriented_ring position arithmetic below is used, unchanged.

    This exists because oriented_ring is a bare ordering and _get_locant_from_oriented_ring can only ever return pos + 1. A fused parent’s numbering is not a position sequence – it skips (4 -> 4a -> 5) and it depends on which of several automorphic numberings the parent name was spelled from. Deriving a substituent locant from position arithmetic while the parent name was spelled from a different numbering names a DIFFERENT molecule, which is exactly what happened to 2-substituted tetralins.

orthonym.assembly.substituent_enumerator.extract_chain_substituents(mol, principal_chain, substituents_dict)#

Convert a substituents dict to a list of SubstituentInfo namedtuples.

Takes the existing features.substituents dict (position -> list of atom index lists) and wraps each entry as a SubstituentInfo for unified downstream processing.

Parameters:
  • mol – RDKit Mol object.

  • principal_chain – List of atom indices in the principal chain.

  • substituents_dict – Dict mapping chain position (1-indexed) to list of substituent atom index lists.

Returns:

List of SubstituentInfo namedtuples.

orthonym.assembly.substituent_enumerator.classify_and_name_fragment(mol, frag_info, parent_atoms, features=None)#

Classify a substituent fragment and produce its IUPAC prefix name.

Routes each fragment to the appropriate naming function based on its composition:

  • fg_only: No carbon atoms (halogens, -OH, -NH2, -NO2, etc.)

  • pure_alkyl: Only carbon atoms (methyl, ethyl, etc.)

  • compound: Carbon + heteroatoms (trifluoromethyl, hydroxymethyl, etc.)

a phase: Tier-0.5 prefix-form check runs FIRST so that any fragment matching the 14-row IUPAC / prefix-form table is named via the canonical prefix form (e.g., -C(=O)OCH3 -> methoxycarbonyl) before falling through to compound/pure_alkyl/fg_only classification. This eliminates the polyfunctional-path duplicate-name bug (hydroxymethyl + methoxycarbonyl on the same ester atoms) per RESEARCH root-cause fix.

Parameters:
  • mol – RDKit Mol object of the full molecule.

  • frag_info – SubstituentInfo namedtuple for this substituent.

  • parent_atoms – Set of atom indices in the parent structure.

  • features – Optional features object (for additional context).

Returns:

IUPAC prefix name string (e.g., “methyl”, “hydroxy”, “trifluoromethyl”), or None if naming fails (with WARNING logged).

orthonym.assembly.substituent_enumerator.collect_substituent_atom_set(substituent_infos)#

Collect the union of all substituent atom indices.

Used by callers to prevent double-counting: FGs whose atoms are entirely within a named branch should be skipped in standalone FG prefix generation.

Parameters:

substituent_infos – List of SubstituentInfo namedtuples.

Returns:

Frozenset of all original mol atom indices covered by substituents.