Batch files#
Put one SMILES per line in a text file. Empty lines are skipped.
$ cat molecules.smi
CCO
c1ccccc1
O=[U](=O)=O
$ orthonym --batch molecules.smi
CCO ethanol
c1ccccc1 benzene
O=[U](=O)=O inorganic compound (not supported)
Each output line is the SMILES, a tab, and the name or the label for a decline. Write the lines to a file with --output, and add --verbose for a count at the end:
$ orthonym --batch molecules.smi --output names.txt --verbose
Processed 3 SMILES, 0 errors
Output written to: names.txt
A line that RDKit cannot read does not stop the run. It gets ERROR: and the reason, and the exit status is then 1:
$ orthonym --batch bad.smi
CCO ethanol
not-a-smiles ERROR: Invalid SMILES: not-a-smiles
CC(=O)O acetic acid
Batch runs use the default tier. --emit-tier and --provenance do not apply to --batch yet. For tiers or provenance rows over many molecules, loop in Python with one Orthonym instance:
import json
from orthonym import Orthonym
namer = Orthonym(general_fallback=True) # the valid tier
with open("molecules.smi") as src, open("rows.jsonl", "w") as out:
for line in src:
smiles = line.strip()
if smiles:
out.write(json.dumps({"smiles": smiles, **namer.name_tiered(smiles)}) + "\n")