orthonym.assembly.per_substring_scoring#
Note
Internal API. Names and behaviour may change between releases.
a phase SCORE-03: per-substring (per-node) scoring on the Name-Tree IR.
Generalizes the scalar ParentCorrectnessScorer (rules/parent_correctness.py)
to NODE granularity (internal notes): for each STRUCTURED NameTreeNode,
attribute a binary parent_score / locant_score / substituent_score
(1.0 match / 0.0 mismatch / 0.5 no-decision) by aligning the node against the
OPSIN-parsed reference name.
- Mechanism (REUSED, not re-invented — RESEARCH §Don’t-Hand-Roll):
The reference name is OPSIN-parsed ONCE per compound (A1 strategy, audit §A1 OPSIN-Cost Prototype) via
opsin_reference_mol; per-node atom alignment is RDKit substructure matching viamatch_token_atoms_in_mol(CanonicalRankAtoms(breakTies=True)deterministic tiebreak — RESEARCH Pitfall 4).Reads the reference name from the SAME thread-local as the parent oracle (
parent_correctness._pc_context). Production (no reference set) ->{}with ZERO OPSIN cost -> byte-identical (SCORE-05).Coarse nodes (
is_coarse_node, single source of truth) are OMITTED from the score dict : they carry no scoreable substrings; the aggregate confidence is retained for them.
POST-HOC contract (, audit §POST-HOC Byte-Identical Contract): the
node_scores dict is attached on CandidateName AFTER compute_confidence
returns; it NEVER feeds back into confidence (contrast the V18
multiple_bond_count recompute at candidate_pool.py:810-823 — the
explicit anti-model).
Binary substituent_score is locked (audit §Binary-vs-Graded); a graded
variant is a documented escape hatch only if the SCORE-04 curated near-tie set
proves binary insufficient (Plan 03 owns that validation).
- class orthonym.assembly.per_substring_scoring.NodeScores(parent_score, locant_score, substituent_score)#
Bases:
objectPer-node parent/locant/substituent scores (binary: 1.0 / 0.0 / 0.5).
- parent_score: float#
- locant_score: float#
- substituent_score: float#
- class orthonym.assembly.per_substring_scoring.PerNodeScorer#
Bases:
objectAttributes per-node parent/locant/substituent scores off the IR .
- static score_tree(tree, mol)#
Return
{id(node): NodeScores}for each STRUCTURED node, or{}.{}is returned when: no reference name is set (production byte-identical guard, zero OPSIN cost), the root is coarse , or the reference name does not OPSIN-parse (no-decision).
- orthonym.assembly.per_substring_scoring.compare_scores(sa, sb)#
Pure lexicographic first-point-of-difference on two NodeScores (internal notes).
a phase STAGE B: extracted from
compare_by_node_scoresso the Stage-B selector can compare explicit keys (e.g. rewritten-tree scores from_scores_from) without re-reading candidate attributes. Behavior on_candidate_scores(a/b)is byte-identical.
- orthonym.assembly.per_substring_scoring.compare_by_node_scores(a, b)#
Lexicographic first-point-of-difference comparator (internal notes).
Compares two near-tie candidates by their ROOT NodeScores in the FIXED priority order
parent_score -> locant_score -> substituent_score. At the first key where they differ, the HIGHER score wins. Returns:-1 a is better +1 b is better
- 0 full per-substring tie -> caller defers to the aggregate
(
select_best_candidate), a STRICT refinement.
This is NOT a weighted sum (the anti-pattern this phase exists to kill — a high
substituent_scoremust never mask a zeroparent_score; the parent must be right before locants matter, mirroring IUPAC first-point-of-difference, Blue Book. Pure function: no OPSIN, no mutation, deterministic.