orthonym.assembly.per_substring_scoring#

Note

Internal API. Names and behaviour may change between releases.

a phase SCORE-03: per-substring (per-node) scoring on the Name-Tree IR.

Generalizes the scalar ParentCorrectnessScorer (rules/parent_correctness.py) to NODE granularity (internal notes): for each STRUCTURED NameTreeNode, attribute a binary parent_score / locant_score / substituent_score (1.0 match / 0.0 mismatch / 0.5 no-decision) by aligning the node against the OPSIN-parsed reference name.

Mechanism (REUSED, not re-invented — RESEARCH §Don’t-Hand-Roll):
  • The reference name is OPSIN-parsed ONCE per compound (A1 strategy, audit §A1 OPSIN-Cost Prototype) via opsin_reference_mol; per-node atom alignment is RDKit substructure matching via match_token_atoms_in_mol (CanonicalRankAtoms(breakTies=True) deterministic tiebreak — RESEARCH Pitfall 4).

  • Reads the reference name from the SAME thread-local as the parent oracle (parent_correctness._pc_context). Production (no reference set) -> {} with ZERO OPSIN cost -> byte-identical (SCORE-05).

  • Coarse nodes (is_coarse_node, single source of truth) are OMITTED from the score dict : they carry no scoreable substrings; the aggregate confidence is retained for them.

POST-HOC contract (, audit §POST-HOC Byte-Identical Contract): the node_scores dict is attached on CandidateName AFTER compute_confidence returns; it NEVER feeds back into confidence (contrast the V18 multiple_bond_count recompute at candidate_pool.py:810-823 — the explicit anti-model).

Binary substituent_score is locked (audit §Binary-vs-Graded); a graded variant is a documented escape hatch only if the SCORE-04 curated near-tie set proves binary insufficient (Plan 03 owns that validation).

class orthonym.assembly.per_substring_scoring.NodeScores(parent_score, locant_score, substituent_score)#

Bases: object

Per-node parent/locant/substituent scores (binary: 1.0 / 0.0 / 0.5).

parent_score: float#
locant_score: float#
substituent_score: float#
class orthonym.assembly.per_substring_scoring.PerNodeScorer#

Bases: object

Attributes per-node parent/locant/substituent scores off the IR .

static score_tree(tree, mol)#

Return {id(node): NodeScores} for each STRUCTURED node, or {}.

{} is returned when: no reference name is set (production byte-identical guard, zero OPSIN cost), the root is coarse , or the reference name does not OPSIN-parse (no-decision).

orthonym.assembly.per_substring_scoring.compare_scores(sa, sb)#

Pure lexicographic first-point-of-difference on two NodeScores (internal notes).

a phase STAGE B: extracted from compare_by_node_scores so the Stage-B selector can compare explicit keys (e.g. rewritten-tree scores from _scores_from) without re-reading candidate attributes. Behavior on _candidate_scores(a/b) is byte-identical.

orthonym.assembly.per_substring_scoring.compare_by_node_scores(a, b)#

Lexicographic first-point-of-difference comparator (internal notes).

Compares two near-tie candidates by their ROOT NodeScores in the FIXED priority order parent_score -> locant_score -> substituent_score. At the first key where they differ, the HIGHER score wins. Returns:

-1 a is better +1 b is better

0 full per-substring tie -> caller defers to the aggregate

(select_best_candidate), a STRICT refinement.

This is NOT a weighted sum (the anti-pattern this phase exists to kill — a high substituent_score must never mask a zero parent_score; the parent must be right before locants matter, mirroring IUPAC first-point-of-difference, Blue Book. Pure function: no OPSIN, no mutation, deterministic.