orthonym.data.polycyclic_data#

Note

Internal API. Names and behaviour may change between releases.

Polycyclic Aromatic Hydrocarbon (PAH) data for IUPAC naming.

Contains lookup tables for common polycyclic aromatic hydrocarbons including: - Canonical SMILES for exact matching - SMARTS patterns for substructure matching - Number of atoms in the PAH core - IUPAC standard numbering (atom index to IUPAC locant mapping) - Allowed substituent positions

IUPAC Naming Rules for PAHs: - PAH numbering is FIXED by IUPAC, not reoriented based on substituents - Use retained names (naphthalene, anthracene, etc.) as parent - Substituents are named with their IUPAC locant position

Reference: IUPAC Blue Book 2013, Section (Fused and Bridged Fused Ring Systems)

PAH Classification: - Bicyclic: naphthalene - Tricyclic: anthracene, phenanthrene, fluorene, acenaphthene, acenaphthylene - Tetracyclic: pyrene, chrysene, tetracene, triphenylene, benz[a]anthracene, benzo[c]phenanthrene - Pentacyclic: pentacene, perylene, benzo[a]pyrene - Hexacyclic+: coronene

Partially saturated PAHs: - 9,10-dihydroanthracene - 1,2-dihydronaphthalene

orthonym.data.polycyclic_data.get_polycyclic_by_smiles(canonical_smiles)#

Look up polycyclic data by canonical SMILES.

Parameters:

canonical_smiles (str) – RDKit canonical SMILES string

Returns:

Dict with ‘name’ and all PAH data if found, None otherwise

Return type:

Dict[str, Any] | None

Example

>>> get_polycyclic_by_smiles('c1ccc2ccccc2c1')
{'name': 'naphthalene', 'canonical_smiles': 'c1ccc2ccccc2c1',...}
orthonym.data.polycyclic_data.get_polycyclic_by_name(name)#

Look up polycyclic data by name.

Parameters:

name (str) – PAH name (e.g., ‘naphthalene’, ‘anthracene’)

Returns:

Dict with PAH data if found, None otherwise

Return type:

Dict[str, Any] | None

orthonym.data.polycyclic_data.is_polycyclic_aromatic(canonical_smiles)#

Check if a canonical SMILES represents a known polycyclic aromatic.

Parameters:

canonical_smiles (str) – RDKit canonical SMILES string

Returns:

True if the SMILES matches a known PAH

Return type:

bool

orthonym.data.polycyclic_data.get_pah_names()#

Get list of all supported PAH names.

Returns:

List of PAH names

Return type:

List[str]

orthonym.data.polycyclic_data.match_polycyclic_core(mol)#

Match a molecule against known PAH cores using substructure matching.

For substituted PAHs, finds the largest matching core and returns the atom mapping from molecule indices to IUPAC locants.

Parameters:

mol – RDKit Mol object

Returns:

Tuple of (pah_name, atom_mapping) where atom_mapping is {mol_idx – iupac_locant} Returns None if no PAH core found

Return type:

Tuple[str, Dict[int, int]] | None

Example

>>> from rdkit import Chem
>>> mol = Chem.MolFromSmiles('Cc1ccc2ccccc2c1') # 2-methylnaphthalene
>>> name, mapping = match_polycyclic_core(mol)
>>> name
'naphthalene'
orthonym.data.polycyclic_data.get_pah_core_atoms(mol, pah_name)#

Get the atom indices that form a PAH core in a molecule.

Parameters:
  • mol – RDKit Mol object

  • pah_name (str) – Name of the PAH (e.g., ‘naphthalene’)

Returns:

Set of atom indices forming the PAH core, or None if no match

Return type:

Set[int] | None

Example

>>> from rdkit import Chem
>>> mol = Chem.MolFromSmiles('Cc1ccc2ccccc2c1') # 2-methylnaphthalene
>>> core = get_pah_core_atoms(mol, 'naphthalene')
>>> len(core) # 10 atoms in naphthalene core
10
orthonym.data.polycyclic_data.get_pah_substituent_positions(mol, pah_name)#

Get IUPAC locant positions where substituents are attached.

Parameters:
  • mol – RDKit Mol object with substituted PAH

  • pah_name (str) – Name of the PAH core (e.g., ‘naphthalene’)

Returns:

List of IUPAC locants with substituents

Return type:

List[int]

Example

>>> from rdkit import Chem
>>> mol = Chem.MolFromSmiles('Cc1ccc2ccccc2c1') # 2-methylnaphthalene
>>> get_pah_substituent_positions(mol, 'naphthalene')
[2] # Methyl at position 2
orthonym.data.polycyclic_data.get_complex_fusion_info(smiles)#

Look up pre-computed fusion descriptor for a known complex polycyclic.

Parameters:

smiles (str) – SMILES string of the compound

Returns:

Tuple of (prefix, descriptor, parent) if found, None otherwise Example: (‘dibenzo’, ‘[a,c]’, ‘anthracene’)

Return type:

Tuple[str, str, str] | None

Example

>>> get_complex_fusion_info('c1ccc2c(c1)cc1ccc3cc4ccccc4cc3c1c2')
('dibenzo', '[a,c]', 'anthracene')
orthonym.data.polycyclic_data.get_complex_fusion_by_name(name)#

Look up complex fusion data by name.

Parameters:

name (str) – Full fusion name (e.g., ‘dibenzo[a,c]anthracene’)

Returns:

Dict with fusion data if found, None otherwise

Return type:

Dict[str, Any] | None

orthonym.data.polycyclic_data.get_edge_letter(ring_name, edge_index)#

Get the IUPAC edge letter for a given edge index in a parent ring.

Parameters:
  • ring_name (str) – Name of the parent ring (e.g., ‘anthracene’)

  • edge_index (int) – 0-indexed edge position

Returns:

Edge letter (‘a’, ‘b’, etc.) or empty string if not found

Return type:

str

Example

>>> get_edge_letter('anthracene', 0)
'a'
>>> get_edge_letter('anthracene', 2)
'c'
orthonym.data.polycyclic_data.get_all_edge_letters(ring_name)#

Get all available edge letters for a parent ring.

Parameters:

ring_name (str) – Name of the parent ring

Returns:

List of edge letters in order

Return type:

List[str]

Example

>>> get_all_edge_letters('benzene')
['a', 'b', 'c', 'd', 'e', 'f']