orthonym.validation.name_morphemes#
Note
Internal API. Names and behaviour may change between releases.
a phase: the independent morpheme-arity oracle.
WHY this module exists#
binding_spine.py proves that a name’s bindings partition the graph (P1),
claim every bond (P2) and account for every charge (P3). All three take the
producer at its word about WHICH atoms a token covers: they check that the
claims are mutually consistent and total, never that a token’s text actually
spells the atoms it claims. A producer that hands the token "methyl" a
six-atom subtree yields a perfectly consistent spine, and P1-P3 pass it.
This module is the independent falsifier for that hole. It answers one
question from the naming data tables alone – “how many heavy atoms does
this token string actually spell?” – with no reference to the binding, the
graph, or the producer that made either. A later proof (P6) compares this
answer against len(binding.atom_ids); a disagreement means the name does
not say what the spine claims it says. Because the two numbers come from
genuinely independent sources (one from the token’s morphemes, one from the
producer’s atom bookkeeping), their agreement is evidence rather than
tautology.
THE SOUNDNESS CONTRACT – the whole value of the module#
A confident answer must be right. A false-confident answer would make P6
reject a CORRECT name, which is the one outcome this milestone cannot ship.
Everything unrecognised, ambiguous, or only partially parsed returns
ArityEstimate(None, False, "<why>"). Coverage is grown later by
measurement, never by guessing.
Four specific refusals follow from that contract, and each is deliberate:
Ambiguity is refusal, not a tie-break. The scanner enumerates EVERY morpheme decomposition of a token rather than committing to the greedy longest match, then requires them all to agree on the total.
"tridecyl"reads astridec|yl(13) or astri|dec|yl(10); a greedy scan would confidently answer one of them, so this module answers neither.A morphotactically impossible reading refuses the WHOLE token – see “WHY AGREEMENT IS NOT ENOUGH” below. This is the repair of the soundness argument itself, not an extra heuristic.
Position-sensitive morphemes are flagged, not hidden.
"ol"is an alcohol suffix (1 O) in suffix position and part of a ring stem elsewhere. Both readings are always enumerated and the suffix reading carries asuffix_onlyflag that the positional rules police; the lexicon is never narrowed bykind, because narrowing it can only manufacture false confidence (again, see below).Whole classes are refused rather than approximated. The lambda (hypervalence) convention and fusion nomenclature are out of scope: fused components SHARE their fusion atoms, so a fusion prefix’s atoms cannot be summed with the base component’s at all.
WHY AGREEMENT IS NOT ENOUGH (the confident-wrong defect)#
“Enumerate every decomposition and require agreement” is sound only if the
correct decomposition is among those enumerated. When a morpheme is MISSING
from the lexicon, a token can still tile – uniquely – into a reading that is
structurally impossible, and unanimity over a set of one then certifies it.
That is not a hypothetical: benzoyl had no benz acyl stem, so its only
tiling was the fusion prefix benzo plus yl, and the oracle answered a
CONFIDENT 6 for an 8-atom acyl group. diazenyl had no diazen parent
hydride, so its only tiling was di|az|en|yl – a multiplier, a skeletal
replacement prefix qualifying nothing, and two bond-order endings – and the
oracle answered a CONFIDENT 0 for two nitrogens. Both would make P6 reject a
Blue Book PIN.
The repair is _well_formed: a small set of morphotactic rules, each read
off an IUPAC construction rule, that say when a tiling cannot describe any
molecule (a replacement prefix that qualifies no skeleton, a suffix that
attaches to no parent hydride, two skeletons juxtaposed with no attachment
affix between them, a fusion prefix, a multiplier governing nothing). A
violation is evidence that the scanner matched across a boundary the lexicon
cannot see – i.e. that the lexicon is INCOMPLETE FOR THIS TOKEN – and once
that is known, “all readings agree” no longer implies “the reading is right”.
So a violation refuses the whole token rather than pruning the offending
reading: pruning would leave exactly the surviving-but-wrong reading that
caused the defect.
WHAT IT REFUSES TO GUESS#
Unknown morphemes, lambda convention, fusion prefixes, von Baeyer/spiro tokens whose bracket descriptor is missing or does not agree with the chain stem beside it, tokens whose multiplied group cannot be delimited, any token admitting two decompositions with different totals, and any token admitting a morphotactically impossible decomposition.
COMPOSITE TOKENS ARE THE COMMON CASE#
Every production binding producer in general_engine.py binds a
substituent’s whole branch under one composed string ("2-chloroethyl",
"4-(methoxymethyl)phenyl"). A single-morpheme oracle would answer
“unverified” for nearly every real binding and P6 would have no teeth, so
token_arity scores a whole token string by summing its morphemes.
A VIEW OVER SHIPPED DATA, NOT A NEW CATALOG#
Every count below is derived from a table this project already ships, so the oracle cannot drift away from the tables the namer itself names from:
Only morphemes that NO shipped table covers are written by hand, in
_STRUCTURAL_AFFIXES below, and each carries a comment saying why it is not
derivable. They are all zero-atom skeletal markers, where “contributes no
atoms” is a definition rather than a measurement.
Being a view over shipped data also means a table entry whose SMILES does not
describe the fragment its name spells would propagate straight through. Four
derivation screens reject exactly that shape, each pointing at a contradiction
inside the entry rather than second-guessing either half of it: a disconnected
SMILES and an all-hydrogen SMILES (both in _heavy_atoms), and a name
asserting an isotope or a Hantzsch-Widman ring the SMILES does not carry (both
in _entry_atoms).
Scope and status: PURE and side-effect-free. token_arity gates nothing and
changes no emitted name – P6 consumes it in audit mode only. (The separate
free_valence_morphology at the foot of the module IS consulted by the
substituent producers; token_arity is not.)
- class orthonym.validation.name_morphemes.ArityEstimate(heavy_atoms, confident, basis)#
Bases:
objectHow many heavy atoms a token spells, or an explicit refusal.
confidentis the only field a caller may branch on for a proof:heavy_atomsis meaningful ONLY whenconfidentis True, and is alwaysNoneotherwise.basisalways explains the answer – which morphemes were read, or why the token was refused – so a measurement pass can group refusals by cause without re-deriving them.- heavy_atoms: int | None#
- confident: bool#
- basis: str#
- orthonym.validation.name_morphemes.lexicon_cache_key()#
identifying the exact inputs of the lexicon build.
- orthonym.validation.name_morphemes.token_arity(token, kind='prefix')#
How many heavy atoms does
tokenspell?kindmay be a:class:~orthonym.validation.binding_spine.BindingKind or the equivalent plain string. It does NOT select a lexicon – there is one lexicon and every reading is always enumerated (see_lexicon). It settles one positional question only: whether a characteristic-group suffix may open the token. For aSUFFIXbinding it may, because the parent hydride it attaches to is a different binding and the token really is just"-oic acid"; for any other kind a leading suffix morpheme is a mis-tiling.A confident answer is guaranteed correct; an unconfident one carries a
basisexplaining the refusal. See the module docstring.
- orthonym.validation.name_morphemes.lexicon_size(include_suffixes=True)#
Number of morphemes the oracle knows. For measurement passes.
include_suffixesis accepted and ignored: there is now one lexicon for every binding kind (see_lexicon), so both answers are the same number.
- orthonym.validation.name_morphemes.lexicon_entries()#
{morpheme: (atoms, category, suffix_only)}– the whole lexicon.Exposed for the arity sweep and the tripwire test: a morpheme that reaches the table without arity data, or a NEW morpheme class nobody has swept against an independent structure, is exactly what this module must not answer confidently about.
- orthonym.validation.name_morphemes.lexicon_sources()#
Names of the derivation steps
_lexiconruns, in order.Pinned by a tripwire test: a NEW morpheme source has not been swept against an independent structure oracle, and its entries have not been checked against
_entry_atoms’ screens, so it must not slip in unnoticed.
- class orthonym.validation.name_morphemes.FreeValenceEstimate(free_valences, confident, basis)#
Bases:
objectHow many free valences a token’s text asserts, or an explicit refusal.
Mirrors:class:ArityEstimate:
confidentis the only field a caller may branch on,free_valencesis meaningful ONLY when it is True and isNoneotherwise, andbasisalways explains the answer.- free_valences: int | None#
- confident: bool#
- basis: str#
- orthonym.validation.name_morphemes.free_valence_morphology(token)#
Read the free-valence count that
token’s own text asserts.