orthonym.metrics.breadth#

Note

Internal API. Names and behaviour may change between releases.

breadth metrics — the pinned definitions behind the breadth instrument.

Why this module exists: every breadth figure quoted before came from an ad-hoc harness in a session a temp dir, and at least one was a ~4x mirage because the harness and the production producer disagreed about what counts as an emission. These functions are the single definition of each metric, unit tested in tests/unit/metrics/test_breadth.py, so a number measured in one session is comparable to a number measured weeks later.

Pure logic: no naming, no OPSIN, no Java. The driver (scripts/measure_breadth.py) supplies per-molecule observations; this module defines what they mean.

orthonym.metrics.breadth.ring_systems(mol)#

Connected ring systems of mol as disjoint atom-index sets.

Two SSSR rings belong to the same system when they share at least one atom, so ortho-fused (naphthalene), bridged and spiro systems each collapse to ONE system, while rings joined only by a bond or a linker (biphenyl) stay separate. This is the ring-SYSTEM grouping; the v1 census classifier that labelled linker-joined rings as ‘fused’ produced a badly wrong topology distribution, so the union-find is deliberate.

orthonym.metrics.breadth.classify_outcome(result)#

EMIT or ABSTAIN for one name_tiered result.

Uses the production failure predicate (errors.is_failure_name) rather than a local truthiness check, because the PIN tier signals abstention with a DESCRIPTIVE string (‘unknown organic compound’) while the best-effort tier signals it with None. Counting the descriptive string as an emission is precisely how a breadth number becomes a mirage.

orthonym.metrics.breadth.parse_refusal_codes(log_lines)#

Distinct refusal codes for ONE molecule, in first-seen order.

Deduplicated on purpose: a code firing five times while the engine probes five substituents is still a single blocker for that molecule, and the milestone question is ‘how many molecules does this site block’, not ‘how often does it fire’. Un-deduplicated counts would overstate the dominant site and mis-rank the build order.

Covers producer refusals (the canonical refusal slugs, general_engine_declined:) AND the two post-hoc gates that suppress a finished name (the self-consistency rejection, opsin_unparseable:) — see _GATE_RES for why omitting the gates made 9 of 11 uncoded abstainers unattributable.

orthonym.metrics.breadth.parse_suppressed_candidates(log_lines)#

The candidate(s) rejected for ONE molecule, from its log lines.

Returns {} when nothing was suppressed, so a caller can row.update it without introducing null keys onto rows the gate never touched.

The LAST suppression is reported as the suppressed candidate because the retry cascade (namer.py:2323-2354) re-submits alternatives and each is re-gated, so the final rejection is the one that ended the molecule. All of them are kept under suppressed_all as well: “one bad candidate” and “the producer kept offering variants of the same wrong structure” need different fixes, and only the full list distinguishes them.

Scope: this reads the line ONLY. The OPSIN validity gate suppresses names too, but for a different reason (unparseable, not wrong molecule), and conflating the two would size a grammar defect as a constitution defect.

⚠ The keys are self01_-prefixed ON PURPOSE, and the prefix is the guard. A molecule can be suppressed by mid-cascade and then terminate at a DIFFERENT gate, so a row whose terminal cause is opsin_unparseable can still carry a payload from earlier in its own cascade. Under the earlier generic names (suppressed_name) that read as “the candidate this row died on”, and an investigator consuming it had to monkeypatch record_suppression to recover the real one. Measured on the 500-row census: 39 of 200 payload-carrying rows terminated somewhere other than self01_mismatch (6 opsin_unparseable, 5 enumerator_last_resort, 1 enumerator_ring_fallback, 27 with no terminal cause recorded). Only trust this payload when ``terminal_detail == ‘self01_mismatch’``.

orthonym.metrics.breadth.residual_refusal_code(row)#

A SPECIFIC named terminal site for an abstainer that logged no code.

Deliberately NOT a catch-all bucket: a single “unattributed” bin would absorb precisely the future instrument gaps this attribution exists to expose, so every branch names its own mechanism and anything unrecognised returns None and is counted as a hard instrument gap.

The three residual mechanisms are structurally unable to produce a log-derived code, which is why they are read out of stored result fields instead:

  • SKIP — RDKit could not parse the input; the namer never ran.

  • EXC — an unhandled exception. NOT a refusal: a crash the instrument’s own except turned into an abstainer row. On the run this was one row, TypeError: '<' not supported between instances of 'str' and 'int'.

  • TIMEOUT — the instrument’s own per-molecule SIGALRM, not an engine decision.

  • LIMIT:<code> — the errors.LIMIT_CATALOG classification carried back in limit_code. The unsupported-element classifier returns before the naming machinery is touched (measured: the [99Tc]-labelled sorbitol row produced ZERO log lines of any kind), so no log stream exists to parse.

Used ONLY where refusal_codes is empty, so it can never displace or inflate a producer-attributed site.

orthonym.metrics.breadth.terminal_site(row)#

The single TERMINAL site for one abstaining row, or None.

Why a terminal basis exists at all — and why it is not merely “the log codes, deduplicated”. Producer refusal-slug / producer_refused:* codes are EXPLORATORY: they fire while the engine searches candidates (five substituent tiers, a ring engine, a chain engine) and are frequently NOT the mechanism that ended the molecule. Measured instance, the class this function exists for:

[I-](CCO)c1ccccc1
  log codes: substituent_is_bare_functional_group:fg_only,
              ring_fragment_declined_by_ring_engine:ring_fragment_declined…,
              n_branch_ring_substituent_unnameable,
              producer_refused:branch unnameable
  terminal: GATE_SUPPRESSED / charge_dropped (namer.py:3916)

A name WAS built for that molecule; a structure-conservation veto removed it. On the log basis four innocent sites collect first/ONLY credit they did not earn, and ONLY is what the build order is ranked by — so the log basis aims a milestone at four sites whose fix would convert nothing. Hence: ONLY and first are computed on TERMINAL attribution; touched stays the UNION of both bases, because the exploratory codes did fire and the multi-blocking depth histogram (the evidence that the census is non-additive) is built from that union.

Three honest caveats, stated because they bound what the terminal basis can be used for:

  1. “Terminal” means the channel’s single attribution, not chronologically last. metrics.abstention is first-writer-wins among generation-stage codes and among post-generation codes, with one documented override: a post-generation site holding a REAL (not failure-marked) candidate overrides a speculative generation-stage record, because the candidate’s existence proves generation completed.

  2. It is coarse — five AbstentionCode values. detail is therefore part of the site key, which splits GATE_SUPPRESSED into the six structure-conservation vetoes plus and the OPSIN validity gate.

  3. Not every post-generation termination is instrumented. The general engine’s inline E1-certificate rejection (namer.py:3586) and its no-jar / dropped-stereo discards (namer.py:3634) throw a generated candidate away without recording, leaving the channel at OTHER with no detail. Those rows get a TERMGAP: site naming the best available log code — never folded into a silent catch-all.

Returns None when the row carries no channel observation at all (an EMIT, a run recorded before the channel was wired, or a crash/timeout in which the namer never returned). Those rows are attributed by the existing residual_refusal_code() machinery instead.

orthonym.metrics.breadth.terminal_stage(site)#

The operational bucket for a terminal site label.

orthonym.metrics.breadth.terminal_basis_available(rows)#

True when rows were measured with the abstention channel wired.

Checked explicitly rather than inferred from “are there any terminal codes”: a run recorded before carries none, and censusing it on the terminal basis would print a clean-looking all-residual zero — the project’s standing “a PERFECT harness result means it did not RUN” failure mode. run_worker stamps terminal_measured on every row it processes, so its ABSENCE is the signal.

orthonym.metrics.breadth.row_attribution(row, basis='log')#

(union, attribution) site lists for one row under basis.

union feeds touched and the depth histogram; attribution feeds first (its head) and ONLY (when it names exactly one site). On the log basis the two are the same list, which is why the log basis credits ONLY to whichever exploratory code happened to fire alone.

orthonym.metrics.breadth.refusal_structure(rows, basis='log')#

Rank refusal sites by ONLY — what a ONE-SITE fix can actually convert.

The touched census cannot size a fix and has already mis-aimed a milestone: per-site touched percentages sum to ~408% of abstainers because a molecule blocked by four sites needs all four cleared, so ring_fragment_declined_by_ring_engine looked like the top target at 116 touched / 95 first-refusals while being the SOLE blocker on exactly 1 molecule. pg='ester' is 12th by touched and the largest single-blocked site at 10.

Three counts per site, all over the SAME denominator (every abstainer, never a filtered subset — a signature computed over a subset is how a lead gets promoted to a defect class):

  • touched — abstainers on which the site fired at all. NOT additive.

  • first — abstainers where it fired first. Systematically over-credits whichever site happens to sit earliest in the pipeline.

  • only — abstainers naming this site and NO other. The honest ceiling of a one-site fix.

single_site_ceiling = (emitted + Σ only) / n and is an OPTIMISTIC UPPER BOUND: len(refusal_codes) == 1 only bounds single-blocked-ness, because an early bail-out hides downstream blockers (project record: “total_refuse=1 does NOT mean single-blocked … 1 of 10 emitted”).

Depth-0 abstainers are reported as uncoded_abstainers and attributed separately via:func:residual_refusal_code. They are NEVER folded into only or the ceiling: uncoded is the opposite of known-single, and counting them would inflate the ceiling with rows whose blocker is unknown.

basis (-T4) selects the ATTRIBUTION source; see row_attribution() and:func:terminal_site for why there are two.

  • "log" (default, unchanged) — every code the engine logged. first and ONLY therefore go to whichever EXPLORATORY producer code fired alone, even when a post-hoc gate is what actually killed the molecule.

  • "terminal" — touched stays the UNION (log ∪ terminal) so the depth histogram keeps its multi-blocking evidence, but first and ONLY are credited to the single terminal mechanism.

⚠ single_site_ceiling is VACUOUS on the terminal basis and the returned ceiling_is_vacuous flag says so. Terminal attribution names exactly one site per row by construction, so every attributed abstainer is “single-blocked” and the ceiling collapses to (emitted + attributed abstainers) / n — a property of the attribution, not a finding about the engine. Use the log-basis ceiling for the multi-blocking bound and the terminal-basis per-site ONLY for the build order; that split is the whole point of carrying both.

orthonym.metrics.breadth.molecule_components(mol)#

Partition heavy atoms into the pieces that must ALL name for the whole molecule to name.

Whole-molecule success is approximately the PRODUCT of per-component success, so this partition is the denominator of per_fragment_p and the reason breadth is multiplicative: at ~3.3 components per molecule, 99% whole-molecule needs p ~= 0.997 per component.

Components are the connected ring systems plus each connected run of acyclic atoms. The partition is total and disjoint by construction.

orthonym.metrics.breadth.aggregate(rows, components_measured=True)#

Roll per-molecule observations into the milestone metrics.

rows entries carry: outcome, tier, refusal_codes, structure_wrong, opsin_unparseable, n_components, n_components_named.

Two separable loss terms are reported, because they map to DIFFERENT build phases and conflating them hides which one is responsible for a flat number:

  • component loss — per_fragment_p < 1: a component cannot be named even standalone (ring / fragment naming gaps).

  • context loss — context_loss: components that name standalone but whose molecule still abstains (the assembly gap).

refusal_structure carries the ONLY-ranked build order (see refusal_structure()); the two flat refusal_census dicts are kept because they are touched counts and existing run records quote them.

projected_emit_independent is E[p**n] over the observed component-count distribution, NOT p**E[n]: the latter is convex in n and understates the independence model materially on any real corpus (8pp on live data). Treat it as an upper bound, not a calibrated predictor — p is measured on components cut out and capped with implicit H, which is easier than naming them in context, so the independence model overshoots observed emit.