Skip to content

New tokenisation #166

Description

@sfluegel05

So far, SMILES tokenisation works by fetch all atom subsection and other characters with a SMILES parser and keeping an updated list.

Problem

This list gets increasingly long. Models can't be used on newer datasets because of unknown tokens. If diverging branches both add tokens, it is difficult to reconcile.

Solution

Fixed-length tokenisaton: Encodes

  • atom type
  • isotope number
  • charge
  • stereoconfiguration
    as one-hot for each atom. This assumes a predetermined list of possible values for each property.

Bonds get a bond type encoding, bond markers get encoded by their number, other symbols remain as-is (assuming there is a fixed number of them).

Activity

  1. sfluegel05 commented on May 6, 2026

    @sfluegel05
    CollaboratorAuthor

    Attempt 1 - Regex + atom decomposition

    A standard approach to SMILES tokenisation seems to be the Regex from Schwaller et al. (2019).

    Changes the standard approach

    • Instead of 1 token per atoms, I decompose atoms into 5 tokens (1 each for element, charge, hydrogens, stereochemistry and isotopes). For all categories (and non-atom tokens), I have a predefined list of possible values. This reduces the number of tokens compared to a product of tokens (which would be roughly 126x13x9x3x250=11 million tokens - unless we spend some effort into researching which combinations of element, charge and isotopes are chemically viable).
    • The regex does not cover dative bonds (-> and <- symbols). This is not standard SMILES notation, but RDKit does it like this, so we should have them as single tokens
    • Wildcards are all the same token. In the SMILES, there are a lot of *, but also [1*], [2*] etc. which are chemically the same, the person drawing them just decided to number them. I added a preprocessing that turns all bracketed tokens with * into *.

    A demo tokeniser implementation:

    from __future__ import annotations
    import re
    from typing import List
    
    from rdkit import Chem
    
    
    def _build_bracket_atoms() -> List[str]:
        """Enumerate chemically meaningful bracketed atoms."""
        pt = Chem.GetPeriodicTable()
        elements = [pt.GetElementSymbol(i) for i in range(1, 119)]
        # aromatic forms used in SMILES
        elements += ["c", "n", "o", "s", "p", "b", "te", "se"]
        charges = range(-5, 8)
        hydrogens = range(9)
        stereo = ["None", "@", "@@"]
        isotopes = range(1, 250) # 228Ra is the heaviest isotope in ChEBI (CHEBI:80505) - leave some space for future additions
    
        tokens = set()
        for el in elements:
            tokens.add(f"element_{el}")
        for ch in charges:
            tokens.add(f"charge_{ch}")
        for h in hydrogens:        
            tokens.add(f"hydrogens_{h}")
        for st in stereo:
            tokens.add(f"stereo_{st}")
        for iso in isotopes:
            tokens.add(f"isotope_{iso}")
        tokens.add("isotope_None")  # for non-isotopic atoms
    
        return list(tokens)
    
    NON_BRACKET_TOKENS = [
        # organic subset elements (unbracketed form)
        #"B", "C", "N", "O", "S", "P", "F", "I", "Cl", "Br",
        #"b", "c", "n", "o", "s", "p",
        # bonds / structure
        "(", ")", "=", "#", "->", "<-", ">>", "-", "+", "/", "\\", ":", ".", "~", "*", "$", "?",
        "@", "@@",
        # ring closures: single-digit
        *[str(d) for d in range(10)],
        # ring closures: %10..%99
        *[f"%{n:02d}" for n in range(10, 100)],
    ]
    
    def _build_default_vocab() -> List[str]:
        """Special tokens + non-bracket symbols + bracketed atoms."""
    
        brackets = _build_bracket_atoms()
    
        # de-duplicate while preserving order (specials first)
        seen, vocab = set(), []
        for tok in NON_BRACKET_TOKENS + brackets:
            if tok not in seen:
                seen.add(tok)
                vocab.append(tok)
        return vocab
    
    
    # change to original regex: added -> and <- for handling dative bonds (we use RDKit-normalised SMILES, this is not a standard SMILES feature)
    SMI_REGEX_PATTERN = r"""(\[[^\]]+]|Br?|Cl?|N|O|S|P|F|I|b|c|n|o|s|p|\(|\)|\.|=|#|->|<-|>>?|-|\+|\\|\/|:|~|@@|@|\?|\*|\$|\%[0-9]{2}|[0-9])"""
    EMBEDDING_OFFSET = 10
    UNKNOWN_TOKEN_IDX = 3
    
    
    class BasicSmilesTokenizer(object):
        """
        Run basic SMILES tokenization using a regex pattern developed by Schwaller et. al.
        This tokenizer is to be used when a tokenizer that does not require the transformers library by HuggingFace is required.
    
        Examples
        --------
        >>> from deepchem.feat.smiles_tokenizer import BasicSmilesTokenizer
        >>> tokenizer = BasicSmilesTokenizer()
        >>> print(tokenizer.tokenize("CC(=O)OC1=CC=CC=C1C(=O)O"))
        ['C', 'C', '(', '=', 'O', ')', 'O', 'C', '1', '=', 'C', 'C', '=', 'C', 'C', '=', 'C', '1', 'C', '(', '=', 'O', ')', 'O']
    
    
        References
        ----------
        .. [1] Philippe Schwaller, Teodoro Laino, Théophile Gaudin, Peter Bolgar, Christopher A. Hunter, Costas Bekas, and Alpha A. Lee
            ACS Central Science 2019 5 (9): Molecular Transformer: A Model for Uncertainty-Calibrated Chemical Reaction Prediction
            1572-1583 DOI: 10.1021/acscentsci.9b00576
        """
    
        def __init__(self, regex_pattern: str = SMI_REGEX_PATTERN):
            """Constructs a BasicSMILESTokenizer.
    
            Parameters
            ----------
            regex: string
                SMILES token regex
            """
            self.regex_pattern = regex_pattern
            self.regex = re.compile(self.regex_pattern)
    
            self.vocab = _build_default_vocab()
            self.vocab_dict = {tok: idx + EMBEDDING_OFFSET for idx, tok in enumerate(self.vocab)}
            self.idx_to_token = {idx: tok for tok, idx in self.vocab_dict.items()}
    
        def _parse_bracket_atom(self, bracket_token: str) -> List[str]:
            """
            Parse a bracketed atom token into 5 components: isotope, element, charge, hydrogens, stereo.
            
            E.g. "[85Kr]" -> ["isotope_85", "element_Kr", "charge_0", "hydrogens_0", "stereo_None"]
            """
            atom_str = bracket_token[1:-1]  # Remove brackets
    
            # special case: any atom containing a * is treated as a wildcard (e.g. [3*:0],[1*])
            if '*' in atom_str:
                return ["*"]
            
            isotope = None
            element = None
            charge = 0
            hydrogens = 0
            stereo = None
            
            pos = 0
            
            # Parse isotope (leading digits)
            iso_str = ""
            while pos < len(atom_str) and atom_str[pos].isdigit():
                iso_str += atom_str[pos]
                pos += 1
            if iso_str:
                isotope = iso_str
            
            # Parse element (1-2 letters)
            if pos < len(atom_str) and (atom_str[pos].isupper() or atom_str[pos].islower()):
                element = atom_str[pos]
                pos += 1
                if pos < len(atom_str) and atom_str[pos].islower():
                    element += atom_str[pos]
                    pos += 1
            
            # Parse stereo (@ or @@)
            if pos < len(atom_str) and atom_str[pos] == '@':
                if pos + 1 < len(atom_str) and atom_str[pos + 1] == '@':
                    stereo = "@@"
                    pos += 2
                else:
                    stereo = "@"
                    pos += 1
            
            # Parse hydrogens (H, H2, H3, H4)
            if pos < len(atom_str) and atom_str[pos] == 'H':
                hydrogens = 1
                pos += 1
                if pos < len(atom_str) and atom_str[pos].isdigit():
                    hydrogens = int(atom_str[pos])
                    pos += 1
            
            # Parse charge (+, -, +2, -2, +3, -3)
            if pos < len(atom_str):
                if atom_str[pos] == '+':
                    pos += 1
                    if pos < len(atom_str) and atom_str[pos].isdigit():
                        charge = int(atom_str[pos])
                        pos += 1
                    else:
                        charge = 1
                elif atom_str[pos] == '-':
                    pos += 1
                    if pos < len(atom_str) and atom_str[pos].isdigit():
                        charge = -int(atom_str[pos])
                        pos += 1
                    else:
                        charge = -1
            
            # Format and return 5 component tokens
            return [
                f"isotope_{isotope if isotope else 'None'}",
                f"element_{element}",
                f"charge_{charge}",
                f"hydrogens_{hydrogens}",
                f"stereo_{stereo if stereo else 'None'}"
            ]
    
        def tokenize(self, text):
            """Tokenize a SMILES string, breaking bracketed atoms into 5 components.
            
            Non-bracketed tokens are returned as-is.
            Bracketed atoms are decomposed into: isotope, element, charge, hydrogens, stereo.
            """
            raw_tokens = [token for token in self.regex.findall(text)]
            tokens = []
            
            for token in raw_tokens:
                if token in NON_BRACKET_TOKENS:
                    tokens.append(token)
                    continue
                if not (token.startswith('[') and token.endswith(']')):
                    token = f"[{token}]"  # Wrap non-bracket tokens in brackets for uniformity
                # Parse bracketed atom into 5 components
                components = self._parse_bracket_atom(token)
                tokens.extend(components)
                
            return tokens
        
        def encode(self, text):
            tokens = self.tokenize(text)
            return [self.vocab_dict.get(token, UNKNOWN_TOKEN_IDX) for token in tokens]
        
        def decode(self, token_ids, skip_special_tokens=False):
            tokens = [self.idx_to_token.get(idx, "[UNK]") for idx in token_ids]
            if skip_special_tokens:
                tokens = [tok for tok in tokens if tok not in self.vocab_dict]
            return "".join(tokens)
    
    # ---- quick self-test ------------------------------------------------------
    if __name__ == "__main__":
        #tok = SmilesTokenizer.build_default()
        #print(f"Vocab size: {tok.vocab_size}")
        tok = BasicSmilesTokenizer()
        print(f"Vocab size: {len(tok.vocab)}")
        examples = [
            "->[se]",
            #"CC(=O)Oc1ccccc1C(=O)O",          # aspirin
            #"C[C@H](N)C(=O)O",                 # L-alanine
            #"[13CH3]CO",                       # isotope
            #"C1CC2(CCCCC2)CC1",                # spiro
            #"c1ccc2c(c1)[nH]cn2",              # benzimidazole with [nH]
        ]
        
        for s in examples:
            tokens = tok.tokenize(s)
            ids = tok.encode(s)
            print(f"\n{s}")
            print(f"  tokens: {tokens}")
            print(f"  ids:    {ids}")
            print(f"  decode: {tok.decode(ids)}")

    Results

    I ran the regex on ChEBI (v251, after removing molecules that can't be processed and normalising SMILES with RDKit) and counted the token frequencies. This is the metric used for coverage with my predefined token list.

    With this tokenizer I get

    Total unique token types:    906                                                                                
    Total token occurrences:     11,740,630
    
    (a) Unique token type coverage (unweighted):
        Covered:   906 / 906  (100.0%)
        Uncovered: 0 / 906  (0.0%)
    
    (b) Token coverage weighted by frequency:
        Covered:   11,740,630 / 11,740,630  (100.0%)
        Uncovered: 0 / 11,740,630  (0.0%)
    
  2. sfluegel05 commented on May 6, 2026

    @sfluegel05
    CollaboratorAuthor

    PubChem evaluation

    I also ran the tokenizer on PubChem (I did not normalise the SMILES for that. This should not change the results however since this is mainly about the occurring atoms.)

    Here are the initial results:

    Total unique token types:    4,459
    Total token occurrences:     6,182,259,538
    
    (a) Unique token type coverage (unweighted):
        Covered:   4,277 / 4,459  (95.9%)
        Uncovered: 182 / 4,459  (4.1%)
    
    (b) Token coverage weighted by frequency:
        Covered:   6,182,257,682 / 6,182,259,538  (100.0%)
        Uncovered: 1,856 / 6,182,259,538  (0.0%)
    
    Uncovered tokens (182), sorted by frequency:
        '[si]'                          count=         939  (0.000%)
        '[siH]'                         count=         361  (0.000%)
        '[Ru+8]'                        count=         171  (0.000%)
        '[si-]'                         count=         129  (0.000%)
        '[si+]'                         count=          40  (0.000%)
        '[Os+8]'                        count=          27  (0.000%)
        '[252Cf]'                       count=           4  (0.000%)
        '[si+2]'                        count=           4  (0.000%)
        '[siH-]'                        count=           4  (0.000%)
        '[253Es]'                       count=           3  (0.000%)
        '[253Cf]'                       count=           2  (0.000%)
        '[siH+]'                        count=           2  (0.000%)
        '[250Cf]'                       count=           1  (0.000%)
        '[250Cm]'                       count=           1  (0.000%)
        '[254Fm]'                       count=           1  (0.000%)
        '[255Fm]'                       count=           1  (0.000%)
        '[257Fm]'                       count=           1  (0.000%)
        '[257Md]'                       count=           1  (0.000%)
        '[250Bk]'                       count=           1  (0.000%)
        '[252Fm]'                       count=           1  (0.000%)
        '[251Cf]'                       count=           1  (0.000%)
        '[254Es]'                       count=           1  (0.000%)
        '[254Cf]'                       count=           1  (0.000%)
        '[250Es]'                       count=           1  (0.000%)
        '[251Es]'                       count=           1  (0.000%)
        '[258Md]'                       count=           1  (0.000%)
        '[253Fm]'                       count=           1  (0.000%)
        '[250Fm]'                       count=           1  (0.000%)
        '[250Md]'                       count=           1  (0.000%)
        '[251Cm]'                       count=           1  (0.000%)
        '[251Bk]'                       count=           1  (0.000%)
        '[251Fm]'                       count=           1  (0.000%)
        '[251Md]'                       count=           1  (0.000%)
        '[251No]'                       count=           1  (0.000%)
        '[252Bk]'                       count=           1  (0.000%)
        '[252Es]'                       count=           1  (0.000%)
        '[252Md]'                       count=           1  (0.000%)
        '[252No]'                       count=           1  (0.000%)
        '[252Lr]'                       count=           1  (0.000%)
        '[253Md]'                       count=           1  (0.000%)
        '[253No]'                       count=           1  (0.000%)
        '[253Lr]'                       count=           1  (0.000%)
        '[253Rf]'                       count=           1  (0.000%)
        '[254Md]'                       count=           1  (0.000%)
        '[254No]'                       count=           1  (0.000%)
        '[254Lr]'                       count=           1  (0.000%)
        '[255Cf]'                       count=           1  (0.000%)
        '[255Es]'                       count=           1  (0.000%)
        '[255Md]'                       count=           1  (0.000%)
        '[255No]'                       count=           1  (0.000%)
        '[255Lr]'                       count=           1  (0.000%)
        '[255Rf]'                       count=           1  (0.000%)
        '[255Db]'                       count=           1  (0.000%)
        '[256Cf]'                       count=           1  (0.000%)
        '[256Es]'                       count=           1  (0.000%)
        '[256Fm]'                       count=           1  (0.000%)
        '[256Md]'                       count=           1  (0.000%)
        '[256No]'                       count=           1  (0.000%)
        '[256Lr]'                       count=           1  (0.000%)
        '[256Rf]'                       count=           1  (0.000%)
        '[256Db]'                       count=           1  (0.000%)
        '[257Es]'                       count=           1  (0.000%)
        '[257No]'                       count=           1  (0.000%)
        '[257Lr]'                       count=           1  (0.000%)
        '[257Rf]'                       count=           1  (0.000%)
        '[257Db]'                       count=           1  (0.000%)
        '[258No]'                       count=           1  (0.000%)
        '[258Lr]'                       count=           1  (0.000%)
        '[258Rf]'                       count=           1  (0.000%)
        '[258Db]'                       count=           1  (0.000%)
        '[258Sg]'                       count=           1  (0.000%)
        '[259Fm]'                       count=           1  (0.000%)
        '[259Md]'                       count=           1  (0.000%)
        '[259No]'                       count=           1  (0.000%)
        '[259Lr]'                       count=           1  (0.000%)
        '[259Rf]'                       count=           1  (0.000%)
        '[259Db]'                       count=           1  (0.000%)
        '[259Sg]'                       count=           1  (0.000%)
        '[260Md]'                       count=           1  (0.000%)
        '[260No]'                       count=           1  (0.000%)
        '[260Lr]'                       count=           1  (0.000%)
        '[260Rf]'                       count=           1  (0.000%)
        '[260Db]'                       count=           1  (0.000%)
        '[260Sg]'                       count=           1  (0.000%)
        '[260Bh]'                       count=           1  (0.000%)
        '[261Lr]'                       count=           1  (0.000%)
        '[261Rf]'                       count=           1  (0.000%)
        '[261Db]'                       count=           1  (0.000%)
        '[261Sg]'                       count=           1  (0.000%)
        '[261Bh]'                       count=           1  (0.000%)
        '[262No]'                       count=           1  (0.000%)
        '[262Lr]'                       count=           1  (0.000%)
        '[262Rf]'                       count=           1  (0.000%)
        '[262Db]'                       count=           1  (0.000%)
        '[262Sg]'                       count=           1  (0.000%)
        '[262Bh]'                       count=           1  (0.000%)
        '[263Rf]'                       count=           1  (0.000%)
        '[263Db]'                       count=           1  (0.000%)
        '[263Sg]'                       count=           1  (0.000%)
        '[264Sg]'                       count=           1  (0.000%)
        '[264Bh]'                       count=           1  (0.000%)
        '[264Hs]'                       count=           1  (0.000%)
        '[265Rf]'                       count=           1  (0.000%)
        '[265Sg]'                       count=           1  (0.000%)
        '[265Bh]'                       count=           1  (0.000%)
        '[265Hs]'                       count=           1  (0.000%)
        '[266Lr]'                       count=           1  (0.000%)
        '[266Db]'                       count=           1  (0.000%)
        '[266Sg]'                       count=           1  (0.000%)
        '[266Bh]'                       count=           1  (0.000%)
        '[266Hs]'                       count=           1  (0.000%)
        '[266Mt]'                       count=           1  (0.000%)
        '[267Rf]'                       count=           1  (0.000%)
        '[267Db]'                       count=           1  (0.000%)
        '[267Sg]'                       count=           1  (0.000%)
        '[267Bh]'                       count=           1  (0.000%)
        '[267Hs]'                       count=           1  (0.000%)
        '[268Db]'                       count=           1  (0.000%)
        '[268Hs]'                       count=           1  (0.000%)
        '[268Mt]'                       count=           1  (0.000%)
        '[269Sg]'                       count=           1  (0.000%)
        '[269Hs]'                       count=           1  (0.000%)
        '[270Db]'                       count=           1  (0.000%)
        '[270Bh]'                       count=           1  (0.000%)
        '[270Hs]'                       count=           1  (0.000%)
        '[270Mt]'                       count=           1  (0.000%)
        '[271Sg]'                       count=           1  (0.000%)
        '[271Bh]'                       count=           1  (0.000%)
        '[271Ds]'                       count=           1  (0.000%)
        '[272Bh]'                       count=           1  (0.000%)
        '[272Rg]'                       count=           1  (0.000%)
        '[273Hs]'                       count=           1  (0.000%)
        '[274Bh]'                       count=           1  (0.000%)
        '[274Mt]'                       count=           1  (0.000%)
        '[274Rg]'                       count=           1  (0.000%)
        '[275Hs]'                       count=           1  (0.000%)
        '[275Mt]'                       count=           1  (0.000%)
        '[276Mt]'                       count=           1  (0.000%)
        '[277Hs]'                       count=           1  (0.000%)
        '[277Mt]'                       count=           1  (0.000%)
        '[277Ds]'                       count=           1  (0.000%)
        '[278Mt]'                       count=           1  (0.000%)
        '[278Rg]'                       count=           1  (0.000%)
        '[278Nh]'                       count=           1  (0.000%)
        '[279Ds]'                       count=           1  (0.000%)
        '[279Rg]'                       count=           1  (0.000%)
        '[280Ds]'                       count=           1  (0.000%)
        '[280Rg]'                       count=           1  (0.000%)
        '[281Ds]'                       count=           1  (0.000%)
        '[281Rg]'                       count=           1  (0.000%)
        '[281Cn]'                       count=           1  (0.000%)
        '[282Ds]'                       count=           1  (0.000%)
        '[282Rg]'                       count=           1  (0.000%)
        '[282Cn]'                       count=           1  (0.000%)
        '[282Nh]'                       count=           1  (0.000%)
        '[283Cn]'                       count=           1  (0.000%)
        '[283Nh]'                       count=           1  (0.000%)
        '[284Cn]'                       count=           1  (0.000%)
        '[284Nh]'                       count=           1  (0.000%)
        '[284Fl]'                       count=           1  (0.000%)
        '[285Cn]'                       count=           1  (0.000%)
        '[285Nh]'                       count=           1  (0.000%)
        '[285Fl]'                       count=           1  (0.000%)
        '[286Cn]'                       count=           1  (0.000%)
        '[286Nh]'                       count=           1  (0.000%)
        '[286Fl]'                       count=           1  (0.000%)
        '[287Fl]'                       count=           1  (0.000%)
        '[287Mc]'                       count=           1  (0.000%)
        '[288Fl]'                       count=           1  (0.000%)
        '[288Mc]'                       count=           1  (0.000%)
        '[289Fl]'                       count=           1  (0.000%)
        '[289Mc]'                       count=           1  (0.000%)
        '[290Nh]'                       count=           1  (0.000%)
        '[290Fl]'                       count=           1  (0.000%)
        '[290Mc]'                       count=           1  (0.000%)
        '[290Lv]'                       count=           1  (0.000%)
        '[291Lv]'                       count=           1  (0.000%)
        '[292Lv]'                       count=           1  (0.000%)
        '[293Lv]'                       count=           1  (0.000%)
        '[293Ts]'                       count=           1  (0.000%)
        '[294Ts]'                       count=           1  (0.000%)
        '[295Og]'                       count=           1  (0.000%)
    

    Accordingly, I

    • added si to the list of aromatic atoms
    • moved the charge limit to +9
    • moved the isotope limit to 300

    After these changes, all tokens can be parsed

  3. linked a pull request that will close this issueStatic Tokenizer #169on May 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions