Repository navigation
New tokenisation #166
Copy link
Copy link
Closed
Description
Activity
Attempt 1 - Regex + atom decomposition
A standard approach to SMILES tokenisation seems to be the Regex from Schwaller et al. (2019).
Changes the standard approach
- Instead of 1 token per atoms, I decompose atoms into 5 tokens (1 each for element, charge, hydrogens, stereochemistry and isotopes). For all categories (and non-atom tokens), I have a predefined list of possible values. This reduces the number of tokens compared to a product of tokens (which would be roughly 126x13x9x3x250=11 million tokens - unless we spend some effort into researching which combinations of element, charge and isotopes are chemically viable).
- The regex does not cover dative bonds (
->and<-symbols). This is not standard SMILES notation, but RDKit does it like this, so we should have them as single tokens - Wildcards are all the same token. In the SMILES, there are a lot of
*, but also[1*],[2*]etc. which are chemically the same, the person drawing them just decided to number them. I added a preprocessing that turns all bracketed tokens with*into*.
A demo tokeniser implementation:
from __future__ import annotations import re from typing import List from rdkit import Chem def _build_bracket_atoms() -> List[str]: """Enumerate chemically meaningful bracketed atoms.""" pt = Chem.GetPeriodicTable() elements = [pt.GetElementSymbol(i) for i in range(1, 119)] # aromatic forms used in SMILES elements += ["c", "n", "o", "s", "p", "b", "te", "se"] charges = range(-5, 8) hydrogens = range(9) stereo = ["None", "@", "@@"] isotopes = range(1, 250) # 228Ra is the heaviest isotope in ChEBI (CHEBI:80505) - leave some space for future additions tokens = set() for el in elements: tokens.add(f"element_{el}") for ch in charges: tokens.add(f"charge_{ch}") for h in hydrogens: tokens.add(f"hydrogens_{h}") for st in stereo: tokens.add(f"stereo_{st}") for iso in isotopes: tokens.add(f"isotope_{iso}") tokens.add("isotope_None") # for non-isotopic atoms return list(tokens) NON_BRACKET_TOKENS = [ # organic subset elements (unbracketed form) #"B", "C", "N", "O", "S", "P", "F", "I", "Cl", "Br", #"b", "c", "n", "o", "s", "p", # bonds / structure "(", ")", "=", "#", "->", "<-", ">>", "-", "+", "/", "\\", ":", ".", "~", "*", "$", "?", "@", "@@", # ring closures: single-digit *[str(d) for d in range(10)], # ring closures: %10..%99 *[f"%{n:02d}" for n in range(10, 100)], ] def _build_default_vocab() -> List[str]: """Special tokens + non-bracket symbols + bracketed atoms.""" brackets = _build_bracket_atoms() # de-duplicate while preserving order (specials first) seen, vocab = set(), [] for tok in NON_BRACKET_TOKENS + brackets: if tok not in seen: seen.add(tok) vocab.append(tok) return vocab # change to original regex: added -> and <- for handling dative bonds (we use RDKit-normalised SMILES, this is not a standard SMILES feature) SMI_REGEX_PATTERN = r"""(\[[^\]]+]|Br?|Cl?|N|O|S|P|F|I|b|c|n|o|s|p|\(|\)|\.|=|#|->|<-|>>?|-|\+|\\|\/|:|~|@@|@|\?|\*|\$|\%[0-9]{2}|[0-9])""" EMBEDDING_OFFSET = 10 UNKNOWN_TOKEN_IDX = 3 class BasicSmilesTokenizer(object): """ Run basic SMILES tokenization using a regex pattern developed by Schwaller et. al. This tokenizer is to be used when a tokenizer that does not require the transformers library by HuggingFace is required. Examples -------- >>> from deepchem.feat.smiles_tokenizer import BasicSmilesTokenizer >>> tokenizer = BasicSmilesTokenizer() >>> print(tokenizer.tokenize("CC(=O)OC1=CC=CC=C1C(=O)O")) ['C', 'C', '(', '=', 'O', ')', 'O', 'C', '1', '=', 'C', 'C', '=', 'C', 'C', '=', 'C', '1', 'C', '(', '=', 'O', ')', 'O'] References ---------- .. [1] Philippe Schwaller, Teodoro Laino, Théophile Gaudin, Peter Bolgar, Christopher A. Hunter, Costas Bekas, and Alpha A. Lee ACS Central Science 2019 5 (9): Molecular Transformer: A Model for Uncertainty-Calibrated Chemical Reaction Prediction 1572-1583 DOI: 10.1021/acscentsci.9b00576 """ def __init__(self, regex_pattern: str = SMI_REGEX_PATTERN): """Constructs a BasicSMILESTokenizer. Parameters ---------- regex: string SMILES token regex """ self.regex_pattern = regex_pattern self.regex = re.compile(self.regex_pattern) self.vocab = _build_default_vocab() self.vocab_dict = {tok: idx + EMBEDDING_OFFSET for idx, tok in enumerate(self.vocab)} self.idx_to_token = {idx: tok for tok, idx in self.vocab_dict.items()} def _parse_bracket_atom(self, bracket_token: str) -> List[str]: """ Parse a bracketed atom token into 5 components: isotope, element, charge, hydrogens, stereo. E.g. "[85Kr]" -> ["isotope_85", "element_Kr", "charge_0", "hydrogens_0", "stereo_None"] """ atom_str = bracket_token[1:-1] # Remove brackets # special case: any atom containing a * is treated as a wildcard (e.g. [3*:0],[1*]) if '*' in atom_str: return ["*"] isotope = None element = None charge = 0 hydrogens = 0 stereo = None pos = 0 # Parse isotope (leading digits) iso_str = "" while pos < len(atom_str) and atom_str[pos].isdigit(): iso_str += atom_str[pos] pos += 1 if iso_str: isotope = iso_str # Parse element (1-2 letters) if pos < len(atom_str) and (atom_str[pos].isupper() or atom_str[pos].islower()): element = atom_str[pos] pos += 1 if pos < len(atom_str) and atom_str[pos].islower(): element += atom_str[pos] pos += 1 # Parse stereo (@ or @@) if pos < len(atom_str) and atom_str[pos] == '@': if pos + 1 < len(atom_str) and atom_str[pos + 1] == '@': stereo = "@@" pos += 2 else: stereo = "@" pos += 1 # Parse hydrogens (H, H2, H3, H4) if pos < len(atom_str) and atom_str[pos] == 'H': hydrogens = 1 pos += 1 if pos < len(atom_str) and atom_str[pos].isdigit(): hydrogens = int(atom_str[pos]) pos += 1 # Parse charge (+, -, +2, -2, +3, -3) if pos < len(atom_str): if atom_str[pos] == '+': pos += 1 if pos < len(atom_str) and atom_str[pos].isdigit(): charge = int(atom_str[pos]) pos += 1 else: charge = 1 elif atom_str[pos] == '-': pos += 1 if pos < len(atom_str) and atom_str[pos].isdigit(): charge = -int(atom_str[pos]) pos += 1 else: charge = -1 # Format and return 5 component tokens return [ f"isotope_{isotope if isotope else 'None'}", f"element_{element}", f"charge_{charge}", f"hydrogens_{hydrogens}", f"stereo_{stereo if stereo else 'None'}" ] def tokenize(self, text): """Tokenize a SMILES string, breaking bracketed atoms into 5 components. Non-bracketed tokens are returned as-is. Bracketed atoms are decomposed into: isotope, element, charge, hydrogens, stereo. """ raw_tokens = [token for token in self.regex.findall(text)] tokens = [] for token in raw_tokens: if token in NON_BRACKET_TOKENS: tokens.append(token) continue if not (token.startswith('[') and token.endswith(']')): token = f"[{token}]" # Wrap non-bracket tokens in brackets for uniformity # Parse bracketed atom into 5 components components = self._parse_bracket_atom(token) tokens.extend(components) return tokens def encode(self, text): tokens = self.tokenize(text) return [self.vocab_dict.get(token, UNKNOWN_TOKEN_IDX) for token in tokens] def decode(self, token_ids, skip_special_tokens=False): tokens = [self.idx_to_token.get(idx, "[UNK]") for idx in token_ids] if skip_special_tokens: tokens = [tok for tok in tokens if tok not in self.vocab_dict] return "".join(tokens) # ---- quick self-test ------------------------------------------------------ if __name__ == "__main__": #tok = SmilesTokenizer.build_default() #print(f"Vocab size: {tok.vocab_size}") tok = BasicSmilesTokenizer() print(f"Vocab size: {len(tok.vocab)}") examples = [ "->[se]", #"CC(=O)Oc1ccccc1C(=O)O", # aspirin #"C[C@H](N)C(=O)O", # L-alanine #"[13CH3]CO", # isotope #"C1CC2(CCCCC2)CC1", # spiro #"c1ccc2c(c1)[nH]cn2", # benzimidazole with [nH] ] for s in examples: tokens = tok.tokenize(s) ids = tok.encode(s) print(f"\n{s}") print(f" tokens: {tokens}") print(f" ids: {ids}") print(f" decode: {tok.decode(ids)}")
Results
I ran the regex on ChEBI (v251, after removing molecules that can't be processed and normalising SMILES with RDKit) and counted the token frequencies. This is the metric used for coverage with my predefined token list.
With this tokenizer I get
Total unique token types: 906 Total token occurrences: 11,740,630 (a) Unique token type coverage (unweighted): Covered: 906 / 906 (100.0%) Uncovered: 0 / 906 (0.0%) (b) Token coverage weighted by frequency: Covered: 11,740,630 / 11,740,630 (100.0%) Uncovered: 0 / 11,740,630 (0.0%)PubChem evaluation
I also ran the tokenizer on PubChem (I did not normalise the SMILES for that. This should not change the results however since this is mainly about the occurring atoms.)
Here are the initial results:
Total unique token types: 4,459 Total token occurrences: 6,182,259,538 (a) Unique token type coverage (unweighted): Covered: 4,277 / 4,459 (95.9%) Uncovered: 182 / 4,459 (4.1%) (b) Token coverage weighted by frequency: Covered: 6,182,257,682 / 6,182,259,538 (100.0%) Uncovered: 1,856 / 6,182,259,538 (0.0%) Uncovered tokens (182), sorted by frequency: '[si]' count= 939 (0.000%) '[siH]' count= 361 (0.000%) '[Ru+8]' count= 171 (0.000%) '[si-]' count= 129 (0.000%) '[si+]' count= 40 (0.000%) '[Os+8]' count= 27 (0.000%) '[252Cf]' count= 4 (0.000%) '[si+2]' count= 4 (0.000%) '[siH-]' count= 4 (0.000%) '[253Es]' count= 3 (0.000%) '[253Cf]' count= 2 (0.000%) '[siH+]' count= 2 (0.000%) '[250Cf]' count= 1 (0.000%) '[250Cm]' count= 1 (0.000%) '[254Fm]' count= 1 (0.000%) '[255Fm]' count= 1 (0.000%) '[257Fm]' count= 1 (0.000%) '[257Md]' count= 1 (0.000%) '[250Bk]' count= 1 (0.000%) '[252Fm]' count= 1 (0.000%) '[251Cf]' count= 1 (0.000%) '[254Es]' count= 1 (0.000%) '[254Cf]' count= 1 (0.000%) '[250Es]' count= 1 (0.000%) '[251Es]' count= 1 (0.000%) '[258Md]' count= 1 (0.000%) '[253Fm]' count= 1 (0.000%) '[250Fm]' count= 1 (0.000%) '[250Md]' count= 1 (0.000%) '[251Cm]' count= 1 (0.000%) '[251Bk]' count= 1 (0.000%) '[251Fm]' count= 1 (0.000%) '[251Md]' count= 1 (0.000%) '[251No]' count= 1 (0.000%) '[252Bk]' count= 1 (0.000%) '[252Es]' count= 1 (0.000%) '[252Md]' count= 1 (0.000%) '[252No]' count= 1 (0.000%) '[252Lr]' count= 1 (0.000%) '[253Md]' count= 1 (0.000%) '[253No]' count= 1 (0.000%) '[253Lr]' count= 1 (0.000%) '[253Rf]' count= 1 (0.000%) '[254Md]' count= 1 (0.000%) '[254No]' count= 1 (0.000%) '[254Lr]' count= 1 (0.000%) '[255Cf]' count= 1 (0.000%) '[255Es]' count= 1 (0.000%) '[255Md]' count= 1 (0.000%) '[255No]' count= 1 (0.000%) '[255Lr]' count= 1 (0.000%) '[255Rf]' count= 1 (0.000%) '[255Db]' count= 1 (0.000%) '[256Cf]' count= 1 (0.000%) '[256Es]' count= 1 (0.000%) '[256Fm]' count= 1 (0.000%) '[256Md]' count= 1 (0.000%) '[256No]' count= 1 (0.000%) '[256Lr]' count= 1 (0.000%) '[256Rf]' count= 1 (0.000%) '[256Db]' count= 1 (0.000%) '[257Es]' count= 1 (0.000%) '[257No]' count= 1 (0.000%) '[257Lr]' count= 1 (0.000%) '[257Rf]' count= 1 (0.000%) '[257Db]' count= 1 (0.000%) '[258No]' count= 1 (0.000%) '[258Lr]' count= 1 (0.000%) '[258Rf]' count= 1 (0.000%) '[258Db]' count= 1 (0.000%) '[258Sg]' count= 1 (0.000%) '[259Fm]' count= 1 (0.000%) '[259Md]' count= 1 (0.000%) '[259No]' count= 1 (0.000%) '[259Lr]' count= 1 (0.000%) '[259Rf]' count= 1 (0.000%) '[259Db]' count= 1 (0.000%) '[259Sg]' count= 1 (0.000%) '[260Md]' count= 1 (0.000%) '[260No]' count= 1 (0.000%) '[260Lr]' count= 1 (0.000%) '[260Rf]' count= 1 (0.000%) '[260Db]' count= 1 (0.000%) '[260Sg]' count= 1 (0.000%) '[260Bh]' count= 1 (0.000%) '[261Lr]' count= 1 (0.000%) '[261Rf]' count= 1 (0.000%) '[261Db]' count= 1 (0.000%) '[261Sg]' count= 1 (0.000%) '[261Bh]' count= 1 (0.000%) '[262No]' count= 1 (0.000%) '[262Lr]' count= 1 (0.000%) '[262Rf]' count= 1 (0.000%) '[262Db]' count= 1 (0.000%) '[262Sg]' count= 1 (0.000%) '[262Bh]' count= 1 (0.000%) '[263Rf]' count= 1 (0.000%) '[263Db]' count= 1 (0.000%) '[263Sg]' count= 1 (0.000%) '[264Sg]' count= 1 (0.000%) '[264Bh]' count= 1 (0.000%) '[264Hs]' count= 1 (0.000%) '[265Rf]' count= 1 (0.000%) '[265Sg]' count= 1 (0.000%) '[265Bh]' count= 1 (0.000%) '[265Hs]' count= 1 (0.000%) '[266Lr]' count= 1 (0.000%) '[266Db]' count= 1 (0.000%) '[266Sg]' count= 1 (0.000%) '[266Bh]' count= 1 (0.000%) '[266Hs]' count= 1 (0.000%) '[266Mt]' count= 1 (0.000%) '[267Rf]' count= 1 (0.000%) '[267Db]' count= 1 (0.000%) '[267Sg]' count= 1 (0.000%) '[267Bh]' count= 1 (0.000%) '[267Hs]' count= 1 (0.000%) '[268Db]' count= 1 (0.000%) '[268Hs]' count= 1 (0.000%) '[268Mt]' count= 1 (0.000%) '[269Sg]' count= 1 (0.000%) '[269Hs]' count= 1 (0.000%) '[270Db]' count= 1 (0.000%) '[270Bh]' count= 1 (0.000%) '[270Hs]' count= 1 (0.000%) '[270Mt]' count= 1 (0.000%) '[271Sg]' count= 1 (0.000%) '[271Bh]' count= 1 (0.000%) '[271Ds]' count= 1 (0.000%) '[272Bh]' count= 1 (0.000%) '[272Rg]' count= 1 (0.000%) '[273Hs]' count= 1 (0.000%) '[274Bh]' count= 1 (0.000%) '[274Mt]' count= 1 (0.000%) '[274Rg]' count= 1 (0.000%) '[275Hs]' count= 1 (0.000%) '[275Mt]' count= 1 (0.000%) '[276Mt]' count= 1 (0.000%) '[277Hs]' count= 1 (0.000%) '[277Mt]' count= 1 (0.000%) '[277Ds]' count= 1 (0.000%) '[278Mt]' count= 1 (0.000%) '[278Rg]' count= 1 (0.000%) '[278Nh]' count= 1 (0.000%) '[279Ds]' count= 1 (0.000%) '[279Rg]' count= 1 (0.000%) '[280Ds]' count= 1 (0.000%) '[280Rg]' count= 1 (0.000%) '[281Ds]' count= 1 (0.000%) '[281Rg]' count= 1 (0.000%) '[281Cn]' count= 1 (0.000%) '[282Ds]' count= 1 (0.000%) '[282Rg]' count= 1 (0.000%) '[282Cn]' count= 1 (0.000%) '[282Nh]' count= 1 (0.000%) '[283Cn]' count= 1 (0.000%) '[283Nh]' count= 1 (0.000%) '[284Cn]' count= 1 (0.000%) '[284Nh]' count= 1 (0.000%) '[284Fl]' count= 1 (0.000%) '[285Cn]' count= 1 (0.000%) '[285Nh]' count= 1 (0.000%) '[285Fl]' count= 1 (0.000%) '[286Cn]' count= 1 (0.000%) '[286Nh]' count= 1 (0.000%) '[286Fl]' count= 1 (0.000%) '[287Fl]' count= 1 (0.000%) '[287Mc]' count= 1 (0.000%) '[288Fl]' count= 1 (0.000%) '[288Mc]' count= 1 (0.000%) '[289Fl]' count= 1 (0.000%) '[289Mc]' count= 1 (0.000%) '[290Nh]' count= 1 (0.000%) '[290Fl]' count= 1 (0.000%) '[290Mc]' count= 1 (0.000%) '[290Lv]' count= 1 (0.000%) '[291Lv]' count= 1 (0.000%) '[292Lv]' count= 1 (0.000%) '[293Lv]' count= 1 (0.000%) '[293Ts]' count= 1 (0.000%) '[294Ts]' count= 1 (0.000%) '[295Og]' count= 1 (0.000%)Accordingly, I
- added
sito the list of aromatic atoms - moved the charge limit to +9
- moved the isotope limit to 300
After these changes, all tokens can be parsed
- added
Metadata
Metadata
Assignees
Labels
No labels
So far, SMILES tokenisation works by fetch all atom subsection and other characters with a SMILES parser and keeping an updated list.
Problem
This list gets increasingly long. Models can't be used on newer datasets because of unknown tokens. If diverging branches both add tokens, it is difficult to reconcile.
Solution
Fixed-length tokenisaton: Encodes
as one-hot for each atom. This assumes a predetermined list of possible values for each property.
Bonds get a bond type encoding, bond markers get encoded by their number, other symbols remain as-is (assuming there is a fixed number of them).