orthonym.assembly.naming_utils#
Note
Internal API. Names and behaviour may change between releases.
Name assembly utility functions for IUPAC name generation.
Pure functions for: - Alkyl substituent naming (methyl through decyl) - Substituent prefix formatting with locants and multipliers - Alphabetization sort keys (IUPAC rules for prefix ordering) - Complex substituent detection and multiplier selection - Vowel elision (terminal ‘e’ removal before vowel suffixes) - Suffix with locants formatting (PIN infix style)
All functions are pure: they accept locants as integers (not atom indices). The mapping from atom indices to locants is handled elsewhere.
- orthonym.assembly.naming_utils.strip_italicized_structural_prefix(name)#
Split a leading italicized structural prefix (
sec-/tert-) off a substituent name.Returns
(remainder, had_prefix). The hyphen of such a prefix is part of a SIMPLE retained name /, BB 16282/16286) and must therefore not be read as evidence that the name is a compound substituent.The remainder is what decides:
tert-butylreduces to the simplebutyland is cited bare, whiletert-butylsulfanylreduces to the still-compoundbutylsulfanyland keeps its enclosing marks.Examples
>>> strip_italicized_structural_prefix("tert-butyl") ('butyl', True) >>> strip_italicized_structural_prefix("sec-butyl") ('butyl', True) >>> strip_italicized_structural_prefix("2-methylpropyl") ('2-methylpropyl', False)
- orthonym.assembly.naming_utils.has_structural_hyphen(name)#
Does
namecarry a hyphen that makes it a COMPOUND substituent?True for every hyphen except a leading italicized structural prefix (b)/(d):
tert-butyl/sec-butylare simple;2-methylpropylandtert-butyl-dimethylsilylare compound).Examples
>>> has_structural_hyphen("tert-butyl") False >>> has_structural_hyphen("2-methylpropyl") True >>> has_structural_hyphen("methyl") False
- orthonym.assembly.naming_utils.multiplier_needs_hyphen(name)#
Does a SIMPLE multiplier joined to
namekeep a hyphen boundary?(b)/(d):
di-tert-butyl(neverditert-butyl),di-sec-butyl. This is the SECOND leg of the italicized-prefix rule — the first (has_structural_hyphen) withholds enclosing marks, this one inserts the hyphen — and it shares the one detection primitive so the two legs can never disagree about what an italicized prefix is.Examples
>>> multiplier_needs_hyphen("tert-butyl") True >>> multiplier_needs_hyphen("methyl") False
- orthonym.assembly.naming_utils.italicized_prefix_is_bare(name)#
Must
namebe cited BARE because its only hyphen is an italicized one?Trueonly whennameleads with an italicized structural prefix AND the remainder is itself simple —tert-butyl/sec-butyl(BB 16286 cites*tert*-butyldi(methyl)phosphane(PIN) with the group bare).Falsefor everything else, including a compound remainder (tert-butylsulfanylkeeps its marks) and any name that does not lead with such a prefix.This is the form a site with its OWN compound test should call: because it returns
Falsefor every non-italicized name, dropping it in front of an existing raw'-' in nametest changes that site’s behaviour for italicized-led names ONLY, and never widens or narrows it for anything else.The remainder is judged by the SAME union
enclose_if_compounduses (needs_bracketsORis_complex_substituent), because neither is complete alone:needs_bracketsmisses the bare two-prefix compoundbutylamino, so using it by itself made this predicate calltert-butylaminobare while the canonical path enclosed it. Erring toward “compound” is the safe direction — it can only keep marks, never drop them.Examples
>>> italicized_prefix_is_bare("tert-butyl") True >>> italicized_prefix_is_bare("tert-butylsulfanyl") False >>> italicized_prefix_is_bare("tert-butylamino") False >>> italicized_prefix_is_bare("2-methylpropyl") False >>> italicized_prefix_is_bare("methyl") False
- orthonym.assembly.naming_utils.strip_alphanumerical_noise(text)#
Remove everything (
the Blue Book) bars from a sort key.Greek letters, isotopic descriptors and stereochemical descriptors are not part of alphanumerical order. THE single definition of that exclusion, so the ~140
alpha_sort_keycall sites cannot each grow their own – one already had (rules/ring_substituents.pyopen-coded_alpha_sort_key(_strip_stereo(nm))).What is removed here does not vanish from the ordering: it reappears at
prefix_citation_sort_keytier 3 via:func:cip_descriptor_rank_key, which is where (j) says configuration belongs.Examples
>>> strip_alphanumerical_noise('(e)-3-phenylprop-2-en-1-yl') '3-phenylprop-2-en-1-yl' >>> strip_alphanumerical_noise('beta-d-glucopyranosyloxy') 'glucopyranosyloxy' >>> strip_alphanumerical_noise('[(2z)-pent-2-en-1-yl]') '[pent-2-en-1-yl]' >>> strip_alphanumerical_noise('[4-2h]benzoyl') 'benzoyl' >>> strip_alphanumerical_noise('decyl') # not a descriptor 'decyl'
- orthonym.assembly.naming_utils.unbranched_alkylidene_name(mol, c_idx, exclude_idx)#
Ylidene name for an UNBRANCHED all-carbon H-saturated chain rooted at the double-bonded carbon c_idx (walking away from exclude_idx): methylidene / ethylidene / propylidene (Wave-2 completion C; shared by the sulfine and azinic-acid namers). Fail-closed None on branching, heteroatoms, rings, charges, or further unsaturation.
- orthonym.assembly.naming_utils.should_omit_locant_one(*, context, chain_length=0, is_ring=False, is_heterocyclic=False, is_monosubstituted=False, fg_type='')#
Determine whether locant-1 should be omitted per IUPAC.
Centralized decision point for ALL locant-1 elision in the pipeline. Every call site that decides whether to omit locant-1 must use this function rather than reimplementing the logic inline.
- Parameters:
context (str) – One of “suffix”, “prefix”, or “bond”.
chain_length (int) – Length of the parent chain (0 for rings).
is_ring (bool) – True if the parent is a ring system.
is_heterocyclic (bool) – True if the ring contains heteroatoms.
is_monosubstituted (bool) – True if only one substituent/FG is present.
fg_type (str) – Functional group type string (e.g., “carboxylic_acid”).
- Returns:
True if locant-1 should be omitted from the name.
- Return type:
bool
- orthonym.assembly.naming_utils.is_only_one_substitutable_position(parent_mol)#
-06 (a phase): topological-symmetry predicate for locant-1 elision.
Returns True iff every substitutable ring position of
parent_molis equivalent — i.e. all H-bearing ring carbons share ONE RDKit canonical rank (CanonicalRankAtoms(breakTies=False)= symmetry equivalence classes). This is the real Blue Book condition “only one kind of substitutable hydrogen” / “no isomer generated by moving”) that replaces thenum_double==1/ ring-monosubstituted COUNT PROXIES, which conflate “one feature” with “ring is symmetric”.parent_molMUST be the PARENT HYDRIDE for the rule being decided, NOT the decorated molecule:substituent-prefix locant: ring + unsaturation, substituents stripped (cyclohexane -> True -> methylcyclohexane; cyclohexene -> False -> 3-bromocyclohex-1-ene; benzene -> True -> toluene).
bond locant (d)/4.4): the SATURATED ring WITH substituents (so an unsubstituted cyclohexene -> cyclohexane -> True -> omit -> ‘cyclohexene’; bromocyclohexene -> bromocyclohexane -> False -> keep the ene-locant).
- orthonym.assembly.naming_utils.get_alkyl_name(carbon_count)#
Get alkyl substituent name for a given carbon count.
Returns ‘methyl’ for 1, ‘ethyl’ for 2,…, ‘decyl’ for 10. For carbon counts > 10, delegates to the centralized chain_names module which supports up to 999 carbons using IUPAC compositional naming.
- Parameters:
carbon_count (int) – Number of carbons in the alkyl chain (1-999).
- Returns:
Alkyl substituent name string.
- Raises:
ValueError – If carbon_count is outside supported range.
- Return type:
str
- orthonym.assembly.naming_utils.simple_multiplier_word(n)#
The simple multiplying prefix for
n, orNonewhen no word can be formed.SIMPLE_MULTIPLIERSis only the BASIC-term table (Table 1.4,,the Blue Book) and stops at20: icosa. Everything above 20 is not tabulated but COMPOSED, per (the Blue Book): the basic terms are cited “in the order opposite to that of the constituent digits in the arabic numbers”, joined without hyphens – 21henicosa, 22docosa, 31hentriaconta(verbatim examples at:2820-:2821).data.chain_names.get_chain_prefixalready implements exactly that composition for chain stems, so the multiplier is its stem plus the terminala; this function is the single place that states the rule.Returns
None– never a bare integer – for anynthe composition cannot spell (n < 2, or outsideget_chain_prefix’s 1-9999 range), so callers FAIL CLOSED instead of emitting a non-word such as"21oxa"or"21cyclo". Noten == 1has no multiplying prefix at all (the unmultiplied form is used), which is why it isNoneand not"": a caller that wants the empty string must say so itself.
- orthonym.assembly.naming_utils.needs_brackets(name)#
Determine if a substituent name is a compound substituent needing parentheses.
Per IUPAC, compound substituents (those that contain locants, hyphens, or functional group prefixes fused with alkyl names) must be enclosed in parentheses when used as prefixes on ring parents.
Simple substituents (single-word names like methyl, chloro, hydroxy) do NOT need parentheses.
Already-bracketed names (starting with ‘(’ or ‘[’) are left alone.
- Parameters:
name (str) – The substituent name (e.g., ‘methyl’, ‘hydroxymethyl’, ‘2-methylpropyl’, ‘(N,N-dimethylamino)’).
- Returns:
True if the substituent needs enclosing parentheses, False otherwise.
- Return type:
bool
Examples
>>> needs_brackets("methyl") False >>> needs_brackets("chloro") False >>> needs_brackets("hydroxy") False >>> needs_brackets("hydroxymethyl") True >>> needs_brackets("carboxymethyl") True >>> needs_brackets("aminoethyl") True >>> needs_brackets("2-methylpropyl") True >>> needs_brackets("(N,N-dimethylamino)") False >>> needs_brackets("methoxy") False >>> needs_brackets("tert-butyl") False >>> needs_brackets("tert-butylsulfanyl") True
- orthonym.assembly.naming_utils.get_bracket_depth(name)#
Determine the current bracket nesting depth of a name.
Returns 0 if no brackets, 1 if contains , 2 if contains , etc. Used to determine what enclosing marks to use at the next level.
- Parameters:
name (str) – A substituent or compound name.
- Returns:
Integer nesting depth (0-3).
- Return type:
int
Examples
>>> get_bracket_depth("methyl") 0 >>> get_bracket_depth("2-methylpropyl") 0 >>> get_bracket_depth("(2-methylpropyl)") 1 >>> get_bracket_depth("2-[(1-methylethyl)]propyl") 2
- orthonym.assembly.naming_utils.compute_nesting_depth(name)#
Compute effective bracket nesting depth per (Dec 2025).
Analyzes a name string and returns the effective nesting depth by counting only nesting-relevant brackets, per the following subsections:
-: Ignore indicated hydrogen parentheses, e.g., (1H), (3H) -: Ignore fusion/spiro/ring assembly/von Baeyer brackets -: Count stereo descriptor and compound locant parentheses -: Escalate if consecutive same-level marks would result -: Isotopic labeling convention (not applicable – not implemented)
- Parameters:
name (str) – The chemical name string to analyze.
- Returns:
Effective nesting depth (0 = no relevant brackets, 1 = has relevant parentheses, 2 = has relevant square brackets, etc.).
- Return type:
int
- orthonym.assembly.naming_utils.apply_enclosing_marks(name, depth=0)#
Apply the IUPAC enclosing marks at the nesting depth.
Nesting order: -> -> { } -> again Depth 0: parentheses Depth 1: square brackets (name already contains parentheses) Depth 2: braces (name already contains brackets)
When depth=-1 (sentinel for auto-detect), calls compute_nesting_depth on the input name to determine the effective starting depth from the name’s existing brackets per (Dec 2025 errata). This makes all callers that use the default depth automatically benefit from subsection-aware nesting.
- Parameters:
name (str) – The compound substituent name (without outer brackets).
depth (int) – Nesting depth (0 = outermost, -1 = auto-detect from name).
- Returns:
Name enclosed in the appropriate bracket type.
- Return type:
str
Examples
>>> apply_enclosing_marks("2-methylpropyl", 0) '(2-methylpropyl)' >>> apply_enclosing_marks("2-methylpropyl", 1) '[2-methylpropyl]' >>> apply_enclosing_marks("2-methylpropyl", 2) '{2-methylpropyl}' >>> apply_enclosing_marks("(R)-butan-2-yl", -1) '[(R)-butan-2-yl]'
- orthonym.assembly.naming_utils.renest_group_and_ancestors(name, open_idx)#
Re-derive the enclosing mark of the group opening at
open_idxand of every group that encloses it, innermost first.For a producer that SPLICES a substituent into an already-marked prefix: the spliced group gets deeper, and every enclosing group must move up the “{[({})]}” order (the Blue Book) with it. Re-marking only the spliced group left ‘3-[[3-(alpha-L-rhamnopyranosyloxy)decanoyl]oxy]’ with a bracket directly inside a bracket. Each group’s mark is what
apply_enclosing_marks(content, -1)gives its content, the same rule every other producer uses, including the consecutive-mark step. Groups outside the chain are left untouched.
- orthonym.assembly.naming_utils.enclose_if_compound(name)#
Enclose a substituent prefix in marks iff it is compound/complex.
A SIMPLE prefix (
methyl,phenyl,cyclohexyl, and the retainedbenzyl) is cited bare — BB2-benzylpyridine(PIN),. A COMPOUND or COMPLEX prefix takes enclosing marks, escalating (-> [ -> { when the name already carries brackets — BB2-[(4-bromophenyl)methyl]pyridine(PIN), /.The compound test is the union of the two existing predicates because neither is complete on its own:
needs_bracketscatches a locant, a hyphen and a fused functional prefix (bromomethyl) but not a bare two-prefix compound;is_complex_substituentcatches the two-prefix compound (cyclohexylmethyl) but not the fused functional prefix.A prefix that is NOT fully enclosed yet CONTAINS an inner enclosing mark is also compound and takes an OUTER (escalating) mark —
cyclohexyl(methyl)amino(an N,N-disubstituted amino core, ->[cyclohexyl(methyl)amino]. Neitherneeds_bracketsnoris_complex_substituentsees this when the core carries no locant or hyphen, so the inner-mark test is the third arm of the union. (The same gap was patched inline at one call site with'(' in name or '[' in name; this centralises it — after the_is_fully_enclosedearly return above, any residual([{is a genuine inner mark, and every leading-mark exception — indicated H(1H), stereo(R)-, fusion[2,3-b]— already carries a digit and is caught byis_complex_substituent, so this arm only newly fires on marks-without-locant.)Idempotent: an already fully-enclosed token is returned untouched.
Examples
>>> enclose_if_compound("benzyl") 'benzyl' >>> enclose_if_compound("cyclohexylmethyl") '(cyclohexylmethyl)' >>> enclose_if_compound("(4-bromophenyl)methyl") '[(4-bromophenyl)methyl]' >>> enclose_if_compound("(2-chloroethyl)") '(2-chloroethyl)' >>> enclose_if_compound("cyclohexyl(methyl)amino") '[cyclohexyl(methyl)amino]'
- orthonym.assembly.naming_utils.is_complex_substituent(name)#
Determine if a substituent name is complex.
A substituent is considered complex if its name contains digits, hyphens, is enclosed in parentheses, or contains embedded substituent multiplier prefixes (e.g., diphenyl, trimethyl within a compound name). Complex substituents require bis/tris/tetrakis multipliers instead of di/tri/tetra, and are enclosed in parentheses in the final name.
Note: Modification prefixes like “tetrahydro-” or “dihydro-” are NOT multipliers and do NOT make a name complex.
- Parameters:
name (str) – The substituent name (e.g., ‘methyl’, ‘1-methylethyl’).
- Returns:
True if the substituent is complex, False otherwise.
- Return type:
bool
Examples
>>> is_complex_substituent("methyl") False >>> is_complex_substituent("ethyl") False >>> is_complex_substituent("1-methylethyl") True >>> is_complex_substituent("2-propyl") True >>> is_complex_substituent("diphenylphosphanyl") True >>> is_complex_substituent("tetrahydropyranyl") False
- orthonym.assembly.naming_utils.suffix_takes_derived_multiplier(suffix_name)#
vs (b): does this SUFFIX take bis/tris, not di/tri?
**P”16.3.3** (:7038): “The basic numerical prefixes ‘di’, ‘tri’, ‘tetra’, etc. are used to indicate a multiplicity of: (a) functional and cumulative suffixes, basic or modified by functional replacement, **with the exception of ‘thioic acid’ and ‘dithioic acid’ described in (b)**”.
So the answer is False for every suffix except that one closed family. A suffix is NEVER decomposed for substituted-ness – doing so read carboxylic acid as carboxy + lic acid and broke 10 gold rows.
Examples
>>> suffix_takes_derived_multiplier("carboxylic acid") False >>> suffix_takes_derived_multiplier("carbodithioic acid") False >>> suffix_takes_derived_multiplier("dithioic acid") True >>> suffix_takes_derived_multiplier("dithioic O-acid") True
- orthonym.assembly.naming_utils.get_suffix_multiplier_prefix(count, suffix_name)#
The multiplier for a FUNCTIONAL or CUMULATIVE SUFFIX (a)).
A separate entry point from:func:get_multiplier_prefix because the two callers are asking different questions and the substituent answer is not the suffix answer. Sharing one function meant a suffix could be handed to a substituent decomposition – the mechanism behind benzene-1,4-biscarboxylic acid.
- orthonym.assembly.naming_utils.is_substituted_substituent(name, *, mol=None, atoms=None, attachment=None)#
Is this substituent prefix SUBSTITUTED, per (c) / (a)?
True-> the derived multipliersbis/tris/tetrakis.False-> the basic multipliersdi/tri/tetra.This answers ONLY the multiplier question. Enclosing marks are decided separately by:func:is_complex_substituent /
needs_brackets(), and the two genuinely disagree:propan-2-ylis enclosed AND simply multiplied —di(propan-2-yl).A prefix is substituted when a DETACHABLE substituent prefix sits in front of its own parent core. A prefix that is nothing but an (unsubstituted) parent hydride with attachment/unsaturation locants is simple, however many digits it carries.
Examples
>>> is_substituted_substituent("propan-2-yl") False >>> is_substituted_substituent("naphthalen-2-yl") False >>> is_substituted_substituent("tridecyl") False >>> is_substituted_substituent("hydroxymethyl") True >>> is_substituted_substituent("2-methylbutyl") True >>> is_substituted_substituent("bromomethyl") True
- orthonym.assembly.naming_utils.needs_p1634_marks(name)#
(c)/(d): does this alkyl prefix take enclosing marks when multiplied?
Returns True for unbranched primary-alkyl names beginning with a numerical- multiplier syllable (decyl C10; dodecyl..nonadecyl C12-C19). Such names are parenthesised when count > 1 — ‘di(dodecyl)silane’ — but keep the basic di/tri multiplier. undecyl (C11) and icosyl+ (C20+) are excluded (no leading multiplier syllable), as are C1-C9. The caller applies the count > 1 gate.
Examples
>>> needs_p1634_marks("dodecyl") True >>> needs_p1634_marks("undecyl") False >>> needs_p1634_marks("methyl") False
- orthonym.assembly.naming_utils.opens_with_replacement_prefix(name)#
(c) (the Blue Book,:7174-7178): does this substituent prefix open, after its locants and indicated hydrogen, with a skeletal replacement (‘a’) prefix, so that ‘di’ in front of it could be read as part of the replacement count (‘di(1,2-oxazol-3-yl)’ as a ‘dioxazole’)? Substituent prefixes (ending in ‘yl’, ‘ylidene’, ‘ylidyne’) and the parent hydrides a multiplicative name multiplies (‘azetidine’ in “1,1’-carbonylbis(azetidine)”, not the ‘diazetidine’ ring) are considered; the acyl prefixes of oxalic and oxamic acid and the oxoacid prefixes (‘phosphono’) are not replacement names.
- orthonym.assembly.naming_utils.get_multiplier_prefix(count, substituent_name)#
Get the appropriate multiplier prefix for a count of substituents.
For count=1, returns empty string (no multiplier needed).
⚠ Docstring corrected 2026-07-30. It said “For simple substituent names (no digits/hyphens), uses di/tri/tetra. For complex substituent names, uses bis/tris/tetrakis.” That is not what this function does, and the difference is load-bearing. The predicate is
is_substituted_substituent(plusCATENATION_AMBIGUOUS_PREFIXES) — notis_complex_substituent, which a trace measured at zero calls from here. “Has digits/hyphens” is the complex test, and the two disagree:1,2-xylene complex=True substituted=False -> ‘di’ propan-2-yl complex=True substituted=False -> ‘di’ bromomethyl complex=True substituted=True -> ‘bis’ sulfanyl complex=False substituted=False -> ‘bis’ (catenation-ambiguous)
The governing rule is ** / **, not (which is alphanumerical order — a section this project has twice cited by mistake for di/bis). The discriminator is a contracted name:
:50981,1-dimethoxypropane (PIN)takesdibecausemethoxyis contracted (:17958), while:353441,1-bis(methylsulfanyl)pentane (PIN)takesbisfor the uncontracted form.⛔ Do not “fix” this by re-pointing at ``is_complex_substituent``. Commit `` did exactly that and regressed in both directions —
bis(propan-2-yl)against the PIN at:25719, anddihydroxymethylagainst (a). Preserve the ordering of the checks.- Parameters:
count (int) – Number of identical substituents.
substituent_name (str) – The substituent name to determine simple vs complex.
- Returns:
Multiplier prefix string, or empty string for count=1.
- Return type:
str
Examples
>>> get_multiplier_prefix(1, "methyl") '' >>> get_multiplier_prefix(2, "methyl") 'di' >>> get_multiplier_prefix(3, "methyl") 'tri' >>> get_multiplier_prefix(2, "1-methylethyl") 'bis' >>> get_multiplier_prefix(3, "1-methylethyl") 'tris'
- orthonym.assembly.naming_utils.multiplied_component(count, name, marked)#
Join the multiplier to a component whose enclosure is already decided.
nameis the BARE prefix name (what decides the multiplier);markedis the same prefix after the caller has applied whatever enclosing marks its own .x context requires. Returns the finished multiplied token.THE one place the multiplier and the hyphen meet, so the ~7 producers that used to keep a private
{1:'', 2:'di', 3:'tri'}table – and therefore could never emitbis/trisat all – cannot drift again.Two rules interact here and both are in:
**** clause (d) (
the Blue Book, verbatimdi-*tert*-butyl): a hyphen separates the multiplier from a BARE italicized prefix.**** (
:6968): “No hyphen is placed after a numerical prefix cited in front of a compound substituent enclosed by parentheses, even if that substituent begins with locants” – so once the caller has enclosed the component, the hyphen is dropped.
Examples
>>> multiplied_component(2, 'methyl', 'methyl') 'dimethyl' >>> multiplied_component(2, 'tert-butyl', 'tert-butyl') 'di-tert-butyl' >>> multiplied_component(2, 'propan-2-yl', '(propan-2-yl)') 'di(propan-2-yl)' >>> multiplied_component(2, 'cyclohexylmethyl', '(cyclohexylmethyl)') 'bis(cyclohexylmethyl)' >>> multiplied_component(2, 'dodecyl', 'dodecyl') 'di(dodecyl)'
- orthonym.assembly.naming_utils.begins_with_italic_designator(name)#
True iff
nameopens with the italic designator of a retained parent (‘s-indacene’, ‘as-indacen-1-yl’). A preceding prefix that ends in a Roman letter is then joined with a hyphen: (the Blue Book, “Hyphens are used in substitutive names:”) (d) “to separate italic letters from Roman letters” (:6960), example ‘as-indacene’ (:6966); ‘…cyclopenta[cd]-s- indacene (PIN)’ (:14505). So ‘3-methyl-s-indacene’, not ‘3-methyls-indacene’.
- orthonym.assembly.naming_utils.starts_with_locant(prefix)#
True iff a formatted substituent prefix opens with a locant set.
'4-chloro'and'N-methyl'and'N,4-dimethyl'-> True;'chloro'and'methyl'-> False.WHY: a locant is always separated from the preceding prefix by a hyphen, and the benzene joiner used to test only
first_char.isdigit. An ITALIC locant starts with a letter, so a merged prefix list glued together as3-chloroN-methylbenzamide– which OPSIN parsed to the correct InChIKey, the documented failure mode where a structural oracle cannot see a spelling defect. Deciding a separator during assembly is not a post-processor; nothing here rewrites a finished name.Fails SAFE: an unrecognised head returns False, i.e. the previous behaviour.
- orthonym.assembly.naming_utils.locant_sort_key(locant)#
Order one locant within a substituent group’s locant set.
ITALIC LETTER locants (
N,N',O,S) sort BEFORE arabic numerals; numerals then sort numerically, and same-letter italics lexically. The Blue Book’s own mixed sets are the authority::42213*N*,*N*,*N*,1-tetramethylquinolin-1-ium-3-aminium (PIN):42460*N*,1,4-triphenyl-1*H*-1,2,4-triazol-4-ium-3-aminide (PIN)
Both show the italic letters leading and ONE multiplying prefix spanning the whole set – which is (
:7038) clause (b) (:7067), “the basic numerical prefixes ‘di’, ‘tri’, ‘tetra’, etc. are used to indicate a multiplicity of:… simple substituent prefixes”, whose own example list printsdimethyl. Multiplicity is a property of the substituent NAME, not of which atom carries it, so anN-methyl and a ring/chainmethylare ONE group of two. “Citation of locants” (:2869) is then deny-by-default and requires the whole locant set to be cited.WHY THIS EXISTS AS A SHARED HELPER: a merged group holds
intandstrlocants together, andsorted(["N", 4])raises TypeError. Every producer that merges an italic bucket into a numeric one needs exactly this key, so it lives here rather than being re-derived per producer.format_substituent_prefixalready renders a merged set correctly once it is ordered – it reproduces both Blue Book examples byte-exactly – so NO new prefix formatter is introduced. (That function’s own docstring records “do NOT add a 4th” predicate; this helper deliberately obeys that.)
- orthonym.assembly.naming_utils.format_substituent_prefix(name, locants, count)#
Format a substituent with locants and multiplier prefix.
name-hygiene fix, item (e) / inventory (a phase — begin; consolidation lands in assembly/parenthesisation fix / a phase). This is the ONE correct substituent-prefix + needs-parens reference /. a phase consolidates the divergent needs-parens / enclosing-mark / prefix-assembly predicates onto THIS function; do NOT add a 4th. The divergent predicates as of 169.7:
rules/amides.py _has_positional_locants(name)
assembly/naming_utils.py is_complex_substituent(name)
assembly/naming_utils.py apply_enclosing_marks(name, depth)
rules/ortho_fused.py format_substituent_prefix(substituents) (different signature)
-FIX Item 9 — this paragraph used to assert three things that are all FALSE today, and a false comment is how the next session acquires a wrong belief, so it is corrected rather than trimmed:
it claimed the predicates “DISAGREE today… on ‘trifluoromethyl’/ ‘tert-butyl’: is_complex_substituent True but _has_positional_locants False”. is_complex_substituent(‘tert-butyl’) is False;
it claimed they are independent. _has_positional_locants now DELEGATES (rules/amides.py: return is_complex_substituent(name)), so the two cannot disagree at all;
it called the consolidation test an “xfail-strict tripwire”. That marker is gone — tests/unit/assembly/test_needs_parens_consolidation.py now plainly asserts the equality and PASSES.
Line numbers are deliberately omitted above: the previous ones (525, 461, 32) had all drifted, which is what made the inventory unfollowable. The consolidation onto THIS function still stands; do NOT add a fourth predicate.
Produces a formatted substituent prefix ready for insertion into an IUPAC name. Handles simple and complex substituents differently: - Simple: locants + multiplier + name (e.g., ‘2,2-dimethyl’) - Complex: locants + multiplier + (name) (e.g., ‘3,5-bis(1-methylethyl)’)
- Parameters:
name (str) – Base substituent name (e.g., ‘methyl’, ‘1-methylethyl’).
locants (List[int]) – List of locant positions for this substituent.
count (int) – Number of identical substituents.
- Returns:
Formatted prefix string.
- Return type:
str
Examples
>>> format_substituent_prefix("methyl", [2], 1) '2-methyl' >>> format_substituent_prefix("methyl", [2, 2], 2) '2,2-dimethyl' >>> format_substituent_prefix("ethyl", [3, 5], 2) '3,5-diethyl' >>> format_substituent_prefix("1-methylethyl", [4, 7], 2) '4,7-bis(1-methylethyl)'
- class orthonym.assembly.naming_utils.AlphaKey(text, original='')#
Bases:
strThe text of a sort key, ORDERED letters first.
### **** ALPHANUMERICAL ORDER(the Blue Book): “Nonitalic Roman letters are considered first… When all the Roman letters are identical, the set of locants… are compared”;****(:3477): “The name of a prefix for a substituent is considered to begin with the first letter of its complete name”. A plain string comparison of the key text lets a hyphen or a digit decide against a letter, because both sort below every lowercase letter in ASCII:'prop-1-en-2-yl'<'propan-2-ylidene'('-'<'a'), although the letters saypropanylidene<propenyl;'hydroxy-4-methylpentyl'<'hydroxymethyl', althoughhydroxymethylis an initial segment ofhydroxymethylpentyland comes first (:21663,4-methyl-3-methylidenehexanoic acid (PIN)). The Blue Book’s own5-(butan-2-yl)-5-butylhentriacontane (PIN)(:3461) is the letter-by-letter reading.The value IS the key text (so
alpha_sort_key('dimethyl') == 'methyl'and every inspection of the text is unchanged); only the ORDER differs fromstr, tier by tier:the Roman letters (
_key_letters) – (:3442), (:3477);the italic structural prefixes of the ORIGINAL name, absence first –
****(:3495) “When… Roman letters do not permit a decision…, italicized letters are considered”: ‘4-butyl-4-tert- butylcyclohexan-1-ol (PIN)’ (:3465). The key text of ‘butyl’ and ‘tert-butyl’ is ‘butyl’ for both, so without this tiersortedkept the input order and one molecule got both spellings;the key text, which carries the locants when the letters tie ,:3517);
the original name – not a rule, it only keeps the order total.
prefix_citation_sort_keytier 1 is the same letters string. Python gives a subclass’s reflected comparison priority, so a plain-string sentinel on the left ('zzzzz' < key) is ordered the same way. Equality with a plain string is string equality (alpha_sort_key('tert-butyl') == 'butyl'); two keys are equal only when every tier is.
- orthonym.assembly.naming_utils.alpha_sort_key(substituent_name)#
alphanumerical sort key for a substituent prefix.
Thin wrapper over:func:_alpha_sort_key_core that removes enclosing marks from the finished key.
⚠ This strip is load-bearing, not cosmetic. The core returns the key with any marks the branch logic left in place, and a leading mark decides comparisons outright because
(is ASCII 40 and[is 91 – both below every lowercase letter.alpha_sort_key('4-[(1R)-1-chloroethyl]phenoxy')used to yield'[1-chloroethyl]phenoxy', which sorts ahead of'chloroethyl'and, through (g), took locant 1 – the wrong numbering for gold row W2F-P8-01. The defect was masked while stereodescriptors were still in the key, since their(sorted earlier still.tests/unit/rules/test_anilino_preferred_prefix.pyhad already recorded the same leak from the(side.The key is only ever COMPARED, never reconstructed into a name, so dropping marks cannot affect any emitted string.
The result is an:class:AlphaKey: its text is the key, its ORDER is ‘s (letters first, then the text), so
'propan-2-ylidene'sorts before'prop-1-en-2-yl'and'hydroxymethyl'before'1-hydroxy-4-methylpentyl'.Examples
>>> alpha_sort_key('4-[(1R)-1-chloroethyl]phenoxy') 'chloroethylphenoxy' >>> alpha_sort_key('(2-chloroethyl)') 'chloroethyl' >>> alpha_sort_key('methyl') 'methyl' >>> alpha_sort_key('propan-2-ylidene') < alpha_sort_key('prop-1-en-2-yl') True
- orthonym.assembly.naming_utils.cip_descriptor_rank_key(prefix)#
rank of a prefix’s configurational symbols, in order.
Lower sorts first, i.e. is cited first and takes the lower locant. Returns an empty tuple when the prefix carries no configuration, so a prefix WITH a descriptor never outranks one without on this tier alone – absence ties with absence and the comparison falls through.
Examples
>>> cip_descriptor_rank_key('(1R)-1-chloroethyl') (0,) >>> cip_descriptor_rank_key('(1S)-1-chloroethyl') (1,) >>> cip_descriptor_rank_key('(2Z)-pent-2-en-1-yl') (0,) >>> cip_descriptor_rank_key('(2E)-pent-2-en-1-yl') (1,) >>> cip_descriptor_rank_key('methyl')
- orthonym.assembly.naming_utils.cip_locant_rank_key(items)#
NUMBERING key of ONE candidate numbering: descriptors WITH locants.
itemsis an iterable of(locant, code)pairs, one per stereodescriptor the name would cite under that numbering (atom_CIPCodevalues R/S/r/s/M/P, bond_CIPCodevalues Z/E).locantmay be any sortable value, as long as one call site passes one kind (ints for rings and chains,(int, str)for natural-product locants such as3aor4'). Lower key = preferred numbering.Why a second key next to:func:cip_descriptor_rank_key: that one ranks the CODES of one prefix in order of appearance – right for citation order , where the locants are already equal. As a NUMBERING key it throws the locants away, so a lone descriptor ties across every numbering and the input atom order decides (
(1r)-,(3r)-and(5r)-1,3,5-trimethyl- cyclohexanefor one molecule, one per SMILES spelling).### **** NUMBERINGclause (j) (the Blue Book): “the lower locant is assigned to CIP stereodescriptors Z, R, M, and r (pseudoasymmetry) that are preferred to E, S, P, and s, respectively”. Two tiers:the
(locant, rank)pairs in locant order, compared at the first point of difference, so a descriptor at a lower locant wins and, at an equal locant, the preferred member of its pair wins. The Blue Book’s own reading of “first point of difference”:(2Z,4S,8R,9E)-undeca-2,9-diene-4,8-diol (PIN)(:3403) “the choice is between ‘E’ and ‘Z’ for position ‘2’, not between ‘R’ and ‘S’ for position ‘4’”. With equal locant sets this orders exactly as the rank tuple of:func:cip_descriptor_rank_key did.only when tier 1 ties exactly: chiral (upper-case) before pseudoasymmetric (lower-case) at the first point of difference – Sequence Rule 4a,
****(:45459) “Chiral stereogenic units precede pseudoasymmetric stereogenic units”; this reproduces hexachlorocyclohexane isomer 2,(1R,2R,3r,4S,5S,6s)(:48131), not(1r,2R,3R,4s,5S,6S).
Returns ```` when no item carries a known descriptor, so stereo-free numberings tie exactly as before.
Examples
>>> cip_locant_rank_key([(1, 'r')]) < cip_locant_rank_key([(3, 'r')]) True >>> cip_locant_rank_key([(2, 'Z'), (4, 'S'), (8, 'R'), (9, 'E')]) < cip_locant_rank_key([(2, 'E'), (4, 'R'), (8, 'S'), (9, 'Z')]) True >>> cip_locant_rank_key()
- orthonym.assembly.naming_utils.prefix_citation_sort_key(prefix, *, parent_locants=False)#
alphanumerical order for citation, as a TOTAL order.
Three tiers, consulted in turn:
the Roman LETTERS only.
### **** ALPHANUMERICAL ORDER’s preamble (the Blue Book) puts “Nonitalic Roman letters… first”, and****(:3448) makes multiplicative prefixes not alter the order already established. Digits and hyphens are dropped, so identical-letter prefixes (pentan-2-ylvspentan-3-yl) tie here on purpose and tier 2 decides.the prefix’s OWN locants, in ORDER OF APPEARANCE.
****(:3517): “When two or more prefixes consist of identical Roman letters, priority for order of citation is given to the group that contains the lowest locant(s) at the first point of difference”, whose first example is this very pair —4-(2-methylbutyl)-N-(3-methylbutyl)- aniline (PIN), “for ordering the substituents ‘2’ is lower than ‘3’” (:3521). Order of appearance rather than a sorted set, because:3533prefers1-(2-methylpentan-3-yl)-1-(3-methylpentan-2-yl)cyclo- pentane (PIN)on the ground that “the locant set ‘2,3’ is lower than ‘3,2’”. Compared vialocant_sort_key, so locant 2 precedes locant 10.the CONFIGURATIONAL symbols, in order of appearance, ranked by
****clause (j) (:3346) /****(:22606).:22589puts this tier exactly here: “since the alphabetic characters and locants (ignoring the configuration symbols) are identical the configurational symbols are compared and ‘R’ precedes ‘S’”. Before this tier existed the descriptor was doing this job from INSIDE tier 1, which(
:3446) forbids – and:44601 says outright that capitalizedCIP descriptors “are written in italics to indicate that they are not involved in the primary stage of alphanumerical order”.
the full string. NOT a nomenclature rule – an engineering requirement. Tiers 1-2 are not injective (
cyclohexylmethylvs a differently-spelled prefix with the same letters and no locants), and a tie in asortedwhose input order came from aCounterover RDKit neighbours resolves to HOW THE SMILES WAS WRITTEN. That is how one molecule got two names:CCC(C)CSSSCCC(C)Cand its own re-spelling gave1-(2-methylbutyl)-3-(3-methylbutyl)trisulfaneand1-(3-methylbutyl)-3-(2-methylbutyl)trisulfane. This tier can only ever be reached once has been exhausted, so it decides nothing the Blue Book decides – it only stops atom order from deciding.
parent_locantsselects which of the TWO input conventions this function is being handed, because the callers genuinely differ and the leading locant means opposite things in them:False(default) – a BARE substituent name (2-methylbutyl,pentan-2-yl). Its leading locant is the substituent’s OWN and is exactly what compares.True– an already-RENDERED prefix string (3,5-dichloro,2-methyl) as produced byformat_substituent_prefix. Its leading locants are PARENT locants, which are assigned BY citation order and therefore may not decide it; they are stripped before tier 2.
Passing the wrong convention was the whole defect: one strip served both, so the bare names lost the locant needs and keyed to
('methylbutyl', )apiece.
- orthonym.assembly.naming_utils.apply_vowel_elision(parent_stem, suffix)#
Apply IUPAC vowel elision rules when joining parent stem and suffix.
The terminal ‘e’ in a parent stem is elided (removed) when the suffix begins with ‘a’, ‘i’, ‘o’, ‘u’, or ‘y’. The ‘e’ is NOT elided before consonants or before another ‘e’.
- Parameters:
parent_stem (str) – Parent name stem (e.g., ‘propane’, ‘butane’, ‘propan’).
suffix (str) – Suffix to append (e.g., ‘ol’, ‘al’, ‘one’, ‘amine’, ‘diol’).
- Returns:
Combined string with elision applied if appropriate.
- Return type:
str
Examples
>>> apply_vowel_elision("propane", "ol") 'propanol' >>> apply_vowel_elision("propane", "al") 'propanal' >>> apply_vowel_elision("propane", "one") 'propanone' >>> apply_vowel_elision("propane", "amine") 'propanamine' >>> apply_vowel_elision("butane", "diol") 'butanediol' >>> apply_vowel_elision("ethane", "oic acid") 'ethanoic acid' >>> apply_vowel_elision("propan", "ol") 'propanol'
- orthonym.assembly.naming_utils.format_suffix_with_locants(parent_stem, unsaturation, suffix, suffix_locants, suffix_multiplier='')#
Assemble a parent name with infix locants in IUPAC 2013 PIN style.
Constructs the parent name by combining stem + unsaturation infix, then attaching the suffix with locants using hyphens. Applies vowel elision rules where appropriate.
- Parameters:
parent_stem (str) – The chain prefix (e.g., ‘prop’, ‘but’, ‘pent’).
unsaturation (str) – Unsaturation infix (e.g., ‘an’, ‘en’, ‘yn’, ‘’).
suffix (str) – The functional group suffix (e.g., ‘ol’, ‘one’, ‘oic acid’, ‘al’).
suffix_locants (List[int]) – Locant positions for the suffix group(s).
suffix_multiplier (str) – Multiplier for multiple suffix groups (e.g., ‘di’, ‘tri’).
- Returns:
Assembled parent name string.
- Return type:
str
Examples
>>> format_suffix_with_locants("prop", "an", "ol", [1]) 'propan-1-ol' >>> format_suffix_with_locants("but", "an", "one", [2]) 'butan-2-one' >>> format_suffix_with_locants("prop", "an", "ol", [1, 2], "di") 'propane-1,2-diol' >>> format_suffix_with_locants("pent", "an", "oic acid", ) 'pentanoic acid' >>> format_suffix_with_locants("prop", "an", "al", ) 'propanal'