orthonym.validation.name_morphemes#

Note

Internal API. Names and behaviour may change between releases.

a phase: the independent morpheme-arity oracle.

WHY this module exists#

binding_spine.py proves that a name’s bindings partition the graph (P1), claim every bond (P2) and account for every charge (P3). All three take the producer at its word about WHICH atoms a token covers: they check that the claims are mutually consistent and total, never that a token’s text actually spells the atoms it claims. A producer that hands the token "methyl" a six-atom subtree yields a perfectly consistent spine, and P1-P3 pass it.

This module is the independent falsifier for that hole. It answers one question from the naming data tables alone – “how many heavy atoms does this token string actually spell?” – with no reference to the binding, the graph, or the producer that made either. A later proof (P6) compares this answer against len(binding.atom_ids); a disagreement means the name does not say what the spine claims it says. Because the two numbers come from genuinely independent sources (one from the token’s morphemes, one from the producer’s atom bookkeeping), their agreement is evidence rather than tautology.

THE SOUNDNESS CONTRACT – the whole value of the module#

A confident answer must be right. A false-confident answer would make P6 reject a CORRECT name, which is the one outcome this milestone cannot ship. Everything unrecognised, ambiguous, or only partially parsed returns ArityEstimate(None, False, "<why>"). Coverage is grown later by measurement, never by guessing.

Four specific refusals follow from that contract, and each is deliberate:

  • Ambiguity is refusal, not a tie-break. The scanner enumerates EVERY morpheme decomposition of a token rather than committing to the greedy longest match, then requires them all to agree on the total. "tridecyl" reads as tridec|yl (13) or as tri|dec|yl (10); a greedy scan would confidently answer one of them, so this module answers neither.

  • A morphotactically impossible reading refuses the WHOLE token – see “WHY AGREEMENT IS NOT ENOUGH” below. This is the repair of the soundness argument itself, not an extra heuristic.

  • Position-sensitive morphemes are flagged, not hidden. "ol" is an alcohol suffix (1 O) in suffix position and part of a ring stem elsewhere. Both readings are always enumerated and the suffix reading carries a suffix_only flag that the positional rules police; the lexicon is never narrowed by kind, because narrowing it can only manufacture false confidence (again, see below).

  • Whole classes are refused rather than approximated. The lambda (hypervalence) convention and fusion nomenclature are out of scope: fused components SHARE their fusion atoms, so a fusion prefix’s atoms cannot be summed with the base component’s at all.

WHY AGREEMENT IS NOT ENOUGH (the confident-wrong defect)#

“Enumerate every decomposition and require agreement” is sound only if the correct decomposition is among those enumerated. When a morpheme is MISSING from the lexicon, a token can still tile – uniquely – into a reading that is structurally impossible, and unanimity over a set of one then certifies it. That is not a hypothetical: benzoyl had no benz acyl stem, so its only tiling was the fusion prefix benzo plus yl, and the oracle answered a CONFIDENT 6 for an 8-atom acyl group. diazenyl had no diazen parent hydride, so its only tiling was di|az|en|yl – a multiplier, a skeletal replacement prefix qualifying nothing, and two bond-order endings – and the oracle answered a CONFIDENT 0 for two nitrogens. Both would make P6 reject a Blue Book PIN.

The repair is _well_formed: a small set of morphotactic rules, each read off an IUPAC construction rule, that say when a tiling cannot describe any molecule (a replacement prefix that qualifies no skeleton, a suffix that attaches to no parent hydride, two skeletons juxtaposed with no attachment affix between them, a fusion prefix, a multiplier governing nothing). A violation is evidence that the scanner matched across a boundary the lexicon cannot see – i.e. that the lexicon is INCOMPLETE FOR THIS TOKEN – and once that is known, “all readings agree” no longer implies “the reading is right”. So a violation refuses the whole token rather than pruning the offending reading: pruning would leave exactly the surviving-but-wrong reading that caused the defect.

WHAT IT REFUSES TO GUESS#

Unknown morphemes, lambda convention, fusion prefixes, von Baeyer/spiro tokens whose bracket descriptor is missing or does not agree with the chain stem beside it, tokens whose multiplied group cannot be delimited, any token admitting two decompositions with different totals, and any token admitting a morphotactically impossible decomposition.

COMPOSITE TOKENS ARE THE COMMON CASE#

Every production binding producer in general_engine.py binds a substituent’s whole branch under one composed string ("2-chloroethyl", "4-(methoxymethyl)phenyl"). A single-morpheme oracle would answer “unverified” for nearly every real binding and P6 would have no teeth, so token_arity scores a whole token string by summing its morphemes.

A VIEW OVER SHIPPED DATA, NOT A NEW CATALOG#

Every count below is derived from a table this project already ships, so the oracle cannot drift away from the tables the namer itself names from:

Only morphemes that NO shipped table covers are written by hand, in _STRUCTURAL_AFFIXES below, and each carries a comment saying why it is not derivable. They are all zero-atom skeletal markers, where “contributes no atoms” is a definition rather than a measurement.

Being a view over shipped data also means a table entry whose SMILES does not describe the fragment its name spells would propagate straight through. Four derivation screens reject exactly that shape, each pointing at a contradiction inside the entry rather than second-guessing either half of it: a disconnected SMILES and an all-hydrogen SMILES (both in _heavy_atoms), and a name asserting an isotope or a Hantzsch-Widman ring the SMILES does not carry (both in _entry_atoms).

Scope and status: PURE and side-effect-free. token_arity gates nothing and changes no emitted name – P6 consumes it in audit mode only. (The separate free_valence_morphology at the foot of the module IS consulted by the substituent producers; token_arity is not.)

class orthonym.validation.name_morphemes.ArityEstimate(heavy_atoms, confident, basis)#

Bases: object

How many heavy atoms a token spells, or an explicit refusal.

confident is the only field a caller may branch on for a proof: heavy_atoms is meaningful ONLY when confident is True, and is always None otherwise. basis always explains the answer – which morphemes were read, or why the token was refused – so a measurement pass can group refusals by cause without re-deriving them.

heavy_atoms: int | None#
confident: bool#
basis: str#
orthonym.validation.name_morphemes.lexicon_cache_key()#

identifying the exact inputs of the lexicon build.

orthonym.validation.name_morphemes.token_arity(token, kind='prefix')#

How many heavy atoms does token spell?

kind may be a:class:~orthonym.validation.binding_spine.BindingKind or the equivalent plain string. It does NOT select a lexicon – there is one lexicon and every reading is always enumerated (see _lexicon). It settles one positional question only: whether a characteristic-group suffix may open the token. For a SUFFIX binding it may, because the parent hydride it attaches to is a different binding and the token really is just "-oic acid"; for any other kind a leading suffix morpheme is a mis-tiling.

A confident answer is guaranteed correct; an unconfident one carries a basis explaining the refusal. See the module docstring.

orthonym.validation.name_morphemes.lexicon_size(include_suffixes=True)#

Number of morphemes the oracle knows. For measurement passes.

include_suffixes is accepted and ignored: there is now one lexicon for every binding kind (see _lexicon), so both answers are the same number.

orthonym.validation.name_morphemes.lexicon_entries()#

{morpheme: (atoms, category, suffix_only)} – the whole lexicon.

Exposed for the arity sweep and the tripwire test: a morpheme that reaches the table without arity data, or a NEW morpheme class nobody has swept against an independent structure, is exactly what this module must not answer confidently about.

orthonym.validation.name_morphemes.lexicon_sources()#

Names of the derivation steps _lexicon runs, in order.

Pinned by a tripwire test: a NEW morpheme source has not been swept against an independent structure oracle, and its entries have not been checked against _entry_atoms’ screens, so it must not slip in unnoticed.

class orthonym.validation.name_morphemes.FreeValenceEstimate(free_valences, confident, basis)#

Bases: object

How many free valences a token’s text asserts, or an explicit refusal.

Mirrors:class:ArityEstimate: confident is the only field a caller may branch on, free_valences is meaningful ONLY when it is True and is None otherwise, and basis always explains the answer.

free_valences: int | None#
confident: bool#
basis: str#
orthonym.validation.name_morphemes.free_valence_morphology(token)#

Read the free-valence count that token’s own text asserts.