orthonym.perception.centres_bridge#

Note

Internal API. Names and behaviour may change between releases.

centres CIP-labelling engine bridge (-03, a phase).

Adopts the centres reference CIP implementation (SiMolecule,.0) as an OPTIONAL source-of-truth for Cahn-Ingold-Prelog stereo descriptors, invoked out-of-process via a batched single-JVM subprocess.

Design : this module mirrors two already-trusted JVM-bridge modules verbatim in posture –

  • validation/opsin_roundtrip.py -> jar resolution + _java_available (5 s probe) + graceful None fallback,

  • validation/dual_validator.py:_parse_batch_with_opsin -> write all SMILES to ONE temp file, a SINGLE subprocess.run (argv list, never a shell), positional <labels>\t<ID> parse, os.unlink in a finally.

No new JVM-bridge logic is invented; the only centres-specific work is the both-endpoint label parsing (centres labels C=C / C=N at BOTH 1-based atom endpoints, e.g. 7E 8E; tetrahedral is a single label, e.g. 2S) and the mapping of those per-atom labels back onto RDKit bond _CIPCode props .

Security posture (threat model T-177-01/02/03):
  • SMILES are passed via the temp FILE, never shell-interpolated, never on argv – subprocess.run([...]) with an explicit argv list, no shell interpretation.

  • The temp file uses tempfile.NamedTemporaryFile (OS-randomised path) and is os.unlink-ed in a finally on EVERY path (incl. timeout / error).

  • The jar is a vendored, unmodified,.0 artifact resolved from PROJECT_ROOT (not from user input); provenance recorded in NOTICE.

This engine is OPT-IN. Callers gate on ORTHONYM_USE_CENTRES_CIP and MUST treat a None return (jar/Java absent) as “fall back to RDKit” – a missing JVM never hard-fails a name .

orthonym.perception.centres_bridge.parse_centres_labels(labels_str)#

Parse a centres label token list into {1-based-atom-idx: descriptor}.

centres emits a space-separated token list per molecule, e.g.:

"2S" -> tetrahedral stereocentre at atom 2 -> {2: 'S'}
"7E 8E" -> C=C labelled at BOTH endpoints -> {7: 'E', 8: 'E'}
"2E 3E" -> C=N labelled at both endpoints -> {2: 'E', 3: 'E'}
"CT4" -> cumulene marker (no atom index) -> ignored here;
               cumulene bond resolution is handled separately.

Tokens are <1-based-int><descriptor> where descriptor is one of R/S/r/s (tetrahedral / pseudoasymmetric), E/Z (cis-trans), M/P/m/p (axial / helical). Tokens without a leading integer (e.g. bare CT4) carry no atom index and are skipped at this layer.

Returns an empty dict for empty / whitespace-only input.

orthonym.perception.centres_bridge.centres_label_batch(smiles_list, timeout=120.0)#

Label a batch of SMILES with centres in a SINGLE JVM invocation.

Mirrors dual_validator._parse_batch_with_opsin: write every SMILES to one temp file (<smiles>\t<position> per row), one subprocess.run (argv list – NO shell), parse the <labels>\t<ID> output by the integer position ID, os.unlink in a finally.

Parameters:
  • smiles_list (List[str]) – SMILES strings to label.

  • timeout (float) – max seconds for the whole batch JVM run.

Returns:

{smiles: {1-based-atom-idx: descriptor}} on success, or None when the jar or Java is absent (caller falls back to RDKit –). A SMILES that centres produced no labels for maps to an empty dict.

Return type:

Dict[str, Dict[int, str]] | None

orthonym.perception.centres_bridge.apply_centres_labels(mol, label_map)#

Set RDKit _CIPCode props on atoms/bonds from a centres label map.

Resolution :
  • Tetrahedral / pseudoasymmetric (R/S/r/s) – single labelled atom: mol.GetAtomWithIdx(idx-1).SetProp('_CIPCode', desc).

  • Cis-trans (E/Z) – centres labels BOTH endpoints of the double bond with the same descriptor; convert each 1-based index to 0-based, find the DOUBLE bond between the two consecutive same-descriptor atoms, and set its _CIPCode. Atoms participating in multiple double bonds (dienes) are handled by matching the specific labelled pair; setting both endpoints’ descriptor onto the one bond is idempotent.

Axial / helical (M/P/m/p) descriptors centres emits for AT/HE compounds are set on the atom directly (RDKit’s downstream consumers read _CIPCode generically). Labels whose target cannot be resolved on this mol are skipped (: a missing descriptor beats a wrong one).

The map uses 1-based atom indices (centres convention). mol is modified in place. This is the production analogue of the validation harness keying.

orthonym.perception.centres_bridge.centres_label_mol(mol)#

Label a single RDKit mol with centres (gated single-mol convenience).

Canonicalises the mol to SMILES, runs the batch path over the one SMILES, and applies the returned labels onto mol’s atoms/bonds via apply_centres_labels.

CRITICAL (/ Phase H): centres labels are 1-based atom indices keyed to the CANONICAL SMILES string emitted by Chem.MolToSmiles, whose atom output order generally DIFFERS from mol’s own atom indices. Applying the labels directly by idx-1 (the pre-Phase-H behaviour) lands descriptors on the WRONG atoms whenever canonicalisation reorders (e.g. 4-hydroxyproline, tartaric acid) – a wrong/malformed descriptor, which Phase H forbids. We remap each centres position through _smilesAtomOutputOrder back to the original atom index before applying. The validation harness (score_suite_centres) is unaffected: it sends the suite’s ORIGINAL SMILES and never round-trips through canonicalisation.

Returns:

True if centres produced a label map and it was applied (the engine was available); False when the caller must fall back to RDKit . False has TWO causes: (1) the engine was unavailable (jar/Java absent), or (2) the engine ran but the _smilesAtomOutputOrder remap could not be read while there were labels to place — declining beats misplacing them. An AVAILABLE engine returning an empty map (an achiral molecule) still returns True: centres ran, there simply were no descriptors.

Return type:

bool