orthonym.assembly.name_comparison#

Note

Internal API. Names and behaviour may change between releases.

/ alphanumerical name comparison + locant ordering.

Pure string logic — no RDKit, no imports from handlers (safe to import from anywhere in assembly/ or rules/ without cycles).

BB (the Blue Book-3195): “Primed locants are placed immediately after the corresponding unprimed locants…; locants consisting of a number and a lower-case letter with or without primes as 4a and 4’a (not 4a’) are placed immediately after the corresponding numeric locant and are followed by locants having superscripts. Italic capital and lower-case letter locants are lower than Greek letter locants, which, in turn, are lower than numerals.”

orthonym.assembly.name_comparison.locant_sort_key(token)#

Total-order key for one locant token per.

Key layout: (class, base, letter, primes, superscript)

class: 0 = italic Roman letter (N/O/S/P), 1 = Greek, 2 = numeral within a numeral: base number, then letter suffix (’’ < ‘a’ < ‘b’), with primes ranking between the bare number and the letter-suffixed forms (4 < 4’ < 4a < 4’a), superscripts last (3a < 3a^1).

The λ bonding-number mark material) does NOT perturb the

order — the base locant decides; the λ value is exposed via

parse_lambda_locant for the comparator (Task 4).

orthonym.assembly.name_comparison.compare_locant_strings(a, b)#
orthonym.assembly.name_comparison.compare_locant_str_sets(set_a, set_b)#

First-point-of-difference over string locant sets.

Mirrors rules/locants.py::compare_locant_sets semantics: both sets are sorted ascending (by locant_sort_key), compared term by term; on a tied prefix the shorter set wins; identical sets return 0.

orthonym.assembly.name_comparison.compare_names(a, b)#

alphanumerical comparison of two complete candidate names.

Returns -1 if a is earlier (preferred as PIN), 1 if b, 0 if equal. Tier 1: Roman letters in order of appearance (italics excluded). Tier 2: italic letters in order of appearance. Tier 3: numerical locants in order of appearance (locant_sort_key each).

orthonym.assembly.name_comparison.parse_lambda_locant(token)#

Split ‘1λ5’ / ‘1lambda5’ / “2’λ4” into (base_locant, bonding_number).

Returns None for tokens without a λ mark.: the λ symbol is ‘cited in conjunction with an appropriate locant’.

orthonym.assembly.name_comparison.compare_lambda_locant_sets(set_a, set_b)#

(the Blue Book): ‘The preferred IUPAC name has the lower locant set for substituent group(s) with the higher bonding number(s) cited as prefixes.’

Only λ-bearing tokens participate; their BASE locants are compared with

first-point-of-difference semantics. Returns 0 when neither

side cites a λ prefix (tier does not apply).