orthonym.data.polycyclic_data#
Note
Internal API. Names and behaviour may change between releases.
Polycyclic Aromatic Hydrocarbon (PAH) data for IUPAC naming.
Contains lookup tables for common polycyclic aromatic hydrocarbons including: - Canonical SMILES for exact matching - SMARTS patterns for substructure matching - Number of atoms in the PAH core - IUPAC standard numbering (atom index to IUPAC locant mapping) - Allowed substituent positions
IUPAC Naming Rules for PAHs: - PAH numbering is FIXED by IUPAC, not reoriented based on substituents - Use retained names (naphthalene, anthracene, etc.) as parent - Substituents are named with their IUPAC locant position
Reference: IUPAC Blue Book 2013, Section (Fused and Bridged Fused Ring Systems)
PAH Classification: - Bicyclic: naphthalene - Tricyclic: anthracene, phenanthrene, fluorene, acenaphthene, acenaphthylene - Tetracyclic: pyrene, chrysene, tetracene, triphenylene, benz[a]anthracene, benzo[c]phenanthrene - Pentacyclic: pentacene, perylene, benzo[a]pyrene - Hexacyclic+: coronene
Partially saturated PAHs: - 9,10-dihydroanthracene - 1,2-dihydronaphthalene
- orthonym.data.polycyclic_data.get_polycyclic_by_smiles(canonical_smiles)#
Look up polycyclic data by canonical SMILES.
- Parameters:
canonical_smiles (str) – RDKit canonical SMILES string
- Returns:
Dict with ‘name’ and all PAH data if found, None otherwise
- Return type:
Dict[str, Any] | None
Example
>>> get_polycyclic_by_smiles('c1ccc2ccccc2c1') {'name': 'naphthalene', 'canonical_smiles': 'c1ccc2ccccc2c1',...}
- orthonym.data.polycyclic_data.get_polycyclic_by_name(name)#
Look up polycyclic data by name.
- Parameters:
name (str) – PAH name (e.g., ‘naphthalene’, ‘anthracene’)
- Returns:
Dict with PAH data if found, None otherwise
- Return type:
Dict[str, Any] | None
- orthonym.data.polycyclic_data.is_polycyclic_aromatic(canonical_smiles)#
Check if a canonical SMILES represents a known polycyclic aromatic.
- Parameters:
canonical_smiles (str) – RDKit canonical SMILES string
- Returns:
True if the SMILES matches a known PAH
- Return type:
bool
- orthonym.data.polycyclic_data.get_pah_names()#
Get list of all supported PAH names.
- Returns:
List of PAH names
- Return type:
List[str]
- orthonym.data.polycyclic_data.match_polycyclic_core(mol)#
Match a molecule against known PAH cores using substructure matching.
For substituted PAHs, finds the largest matching core and returns the atom mapping from molecule indices to IUPAC locants.
- Parameters:
mol – RDKit Mol object
- Returns:
Tuple of (pah_name, atom_mapping) where atom_mapping is {mol_idx – iupac_locant} Returns None if no PAH core found
- Return type:
Tuple[str, Dict[int, int]] | None
Example
>>> from rdkit import Chem >>> mol = Chem.MolFromSmiles('Cc1ccc2ccccc2c1') # 2-methylnaphthalene >>> name, mapping = match_polycyclic_core(mol) >>> name 'naphthalene'
- orthonym.data.polycyclic_data.get_pah_core_atoms(mol, pah_name)#
Get the atom indices that form a PAH core in a molecule.
- Parameters:
mol – RDKit Mol object
pah_name (str) – Name of the PAH (e.g., ‘naphthalene’)
- Returns:
Set of atom indices forming the PAH core, or None if no match
- Return type:
Set[int] | None
Example
>>> from rdkit import Chem >>> mol = Chem.MolFromSmiles('Cc1ccc2ccccc2c1') # 2-methylnaphthalene >>> core = get_pah_core_atoms(mol, 'naphthalene') >>> len(core) # 10 atoms in naphthalene core 10
- orthonym.data.polycyclic_data.get_pah_substituent_positions(mol, pah_name)#
Get IUPAC locant positions where substituents are attached.
- Parameters:
mol – RDKit Mol object with substituted PAH
pah_name (str) – Name of the PAH core (e.g., ‘naphthalene’)
- Returns:
List of IUPAC locants with substituents
- Return type:
List[int]
Example
>>> from rdkit import Chem >>> mol = Chem.MolFromSmiles('Cc1ccc2ccccc2c1') # 2-methylnaphthalene >>> get_pah_substituent_positions(mol, 'naphthalene') [2] # Methyl at position 2
- orthonym.data.polycyclic_data.get_complex_fusion_info(smiles)#
Look up pre-computed fusion descriptor for a known complex polycyclic.
- Parameters:
smiles (str) – SMILES string of the compound
- Returns:
Tuple of (prefix, descriptor, parent) if found, None otherwise Example: (‘dibenzo’, ‘[a,c]’, ‘anthracene’)
- Return type:
Tuple[str, str, str] | None
Example
>>> get_complex_fusion_info('c1ccc2c(c1)cc1ccc3cc4ccccc4cc3c1c2') ('dibenzo', '[a,c]', 'anthracene')
- orthonym.data.polycyclic_data.get_complex_fusion_by_name(name)#
Look up complex fusion data by name.
- Parameters:
name (str) – Full fusion name (e.g., ‘dibenzo[a,c]anthracene’)
- Returns:
Dict with fusion data if found, None otherwise
- Return type:
Dict[str, Any] | None
- orthonym.data.polycyclic_data.get_edge_letter(ring_name, edge_index)#
Get the IUPAC edge letter for a given edge index in a parent ring.
- Parameters:
ring_name (str) – Name of the parent ring (e.g., ‘anthracene’)
edge_index (int) – 0-indexed edge position
- Returns:
Edge letter (‘a’, ‘b’, etc.) or empty string if not found
- Return type:
str
Example
>>> get_edge_letter('anthracene', 0) 'a' >>> get_edge_letter('anthracene', 2) 'c'
- orthonym.data.polycyclic_data.get_all_edge_letters(ring_name)#
Get all available edge letters for a parent ring.
- Parameters:
ring_name (str) – Name of the parent ring
- Returns:
List of edge letters in order
- Return type:
List[str]
Example
>>> get_all_edge_letters('benzene') ['a', 'b', 'c', 'd', 'e', 'f']