Filtering Theory

Introduction

Filtering attempts to eliminate inappropriate or undesirable compounds from a large set before beginning to use them in modeling studies. The goal is to remove all of the compounds that should not be suggested to a medicinal chemist as a potential hit. This exercise is obviously case dependent, depending on ease of the assay, intended target, needs of the user, and so on.

To match this need, the default filter encapsulates many of the standard filtering principles, such as removal of unstable, reactive, and toxic moieties. In addition, FILTER allows the customization of the filtering criteria to fit specific needs.

The criteria for passing or failing a given molecule fall into three categories.

  • Physical properties

    • Molecular weight

    • Topological polar surface area (TPSA)

    • logP

    • Bioavailability

  • Atomic and functional group content

    • Absolute and relative content of heteroatoms

    • Limits on a very wide variety of functional groups

  • Molecular graph topology

    • Number and size of ring systems

    • Flexibility of the molecule

    • Size and shape of non-ring chains

All of the data generated in filtering molecules can be written to a tab-separated file for easy import into a spreadsheet. This functionality allows for combining the values dynamically for a variety of purposes, including, but not limited to, determining which filter values best fit each project’s needs.

History

When OpenEye’s work on filtering technology began in 2000, it was designed simply to remove compounds with reactive or otherwise undesirable functional groups. Over the years, the understanding of lead-like and drug-like compound selection has advanced. In addition, with the publication of Lipinski’s “Rule of 5” [Lipinski-1997], more and more pharmacokinetic properties have appeared earlier in the virtual screening process.

In addition to providing basic functional group selection, the technology is a one-stop database preparation tool aimed at generating databases suitable for high-throughput virtual screening. The following are some components of this tool.

  • Cheminformatics quality control

    • Valence-state validation

    • Aromaticity perception

    • Implicit hydrogen perception

    • Bond-order perception

  • Database preparation

    • Setting pKa states

    • Applying normalizations (tautomers and dative or hypervalent states)

  • Compound selection

Finally, it should be pointed out that in the virtual screening world, time is of the essence. Algorithms for preliminary database preparation should not take large amounts of time. Because of this, all the calculations included are 2D or graph-based algorithms. While this does occasionally limit the technology, it allows for the delivery of a product that is appropriate for the task of virtual-screening database preparation.

A Detailed Examination of False Positives and False Negatives

Nearly every computational tool used in early drug discovery yields statistically predictive, rather than absolutely definitive results. In nearly every case, prudence demands that one consider the causes of false positives and false negatives and make an attempt to optimize the area under the receiver-operator curve (ROC) for the computational tool. However, there are well-known methods for improving statistical predictions of this nature that are independent of the absolute false-positive and false-negative rates. These methods include filtering the population to which a test will be applied. By applying a test to smaller populations that only contain molecules appropriate for the specific application at hand, the negative impact of the false-positive rate on the predictive results can be dramatically improved.

A familiar example from the medical world will serve to illustrate this principle. Assume we have a test for the presence of the new foo virus which has an exceptional ROC curve with false-positive and false-negative values (1/1,000 and 1/1,000 respectively). Let us assume that the foo syndrome, caused by the foo virus, affects 1 person in 20,000. If we gave this test to 100,000 people from the general population, we would expect 5 to actually have the foo syndrome. With this test, there is only a 0.05% percent chance that any of them would not be detected (that is, be a false negative). However, we would expect there to be 100 false-positive test results. Thus of the 105 total positive test results, only 4.8% would actually have the foo syndrome (positive predictive value = 4.8%).

Confusion table for the unfiltered foo virus test (prevalence 1 in 20,000)

Actual Positive

Actual Negative

Predicted Positive

True Positives = 5

False Positives = 100

Positive Predictive Value = 4.8%

Predicted Negative

False Negatives = 0

True Negatives = 99,895

Alternatively, we could start by using very simple screening before applying the test. We first eliminate people who do not have any risk factors for contracting the foo virus. Next, we may eliminate people whose blood is incompatible with the test for the foo virus. Further, we may want to eliminate people who acknowledge that they will refuse treatment for the foo virus even if we determine that they do have it. By these admittedly simple screens, we apply the test for the foo virus to a much smaller group with a decidedly higher prevalence of the virus. For instance, after the filtering, we may be left with a group of only 1,000 people who have a 1 in 200 chance of having the syndrome. Now, we still have the same five people who actually have the disease, but we only expect one false-positive test. Now there are six total positive tests, and 83% of them actually have the syndrome! This is reflected in a much more reasonable (83%) positive predictive value.

Confusion table for the filtered foo virus test (prevalence 1 in 200)

Actual Positive

Actual Negative

Predicted Positive

True Positives = 5

False Positives = 1

Positive Predictive Value = 83%

Predicted Negative

False Negatives = 0

True Negatives = 994

Returning to drug design, in a ligand-based design tool such as ROCS, we can posit that the receiver-operator curve may have a false positive rate as low as 1 in 10,000. For this exercise, assume no false negatives. When using that to identify 50 inhibitors from a database of 2.5 million available compounds, we would identify 300 potential inhibitors, and 5 out of every 6 of these would be a false positive (positive predictive value of 17%). However, running filter first would eliminate 65% of the 2.5 million compounds, leaving 875,000 compounds to run through ROCS. There will be approximately 88 false positives along with the 50 true positives increasing the positive predictive value more than two-fold with relatively little work.

Confusion table for the unfiltered ROCS virtual screen

Actual Positive

Actual Negative

Predicted Positive

True Positives = 50

False Positives = 250

Positive Predictive Value = 17%

Predicted Negative

False Negatives = 0

True Negatives = 2,499,700

Confusion table for the filtered ROCS virtual screen

Actual Positive

Actual Negative

Predicted Positive

True Positives = 50

False Positives = 88

Positive Predictive Value = 36%

Predicted Negative

False Negatives = 0

True Negatives = 874,862

Filtering Principles

The same principle of increasing the positive predictive value by removing obvious true negatives applies to screening for lead candidates, regardless of whether it is virtual screening or high-throughput screening (HTS). While both are reasonable screens, each can be plagued by very low positive predictive values (despite low false-positive rates), particularly when applied to all available compounds or large virtual libraries. Simple filtering techniques focus the set of compounds passed on to more computationally intensive screening methods.

The first approach is to filter based on functional groups. Generally speaking, there are toxic and reactive functional groups that should be removed from consideration, such as alkyl bromides or metals. There are also functional groups that are not strictly forbidden, but are undesirable in large quantities. For instance, parafluorobenzene or trifluoromethyl have specific purposes, but heavily fluorinated molecules can be eliminated.

Beyond simple functional group filtering, both simple and complex physical properties can be used to characterize the kinds of compounds to keep or to eliminate. These properties attempt to consider drug-like qualities, such as bioavailability, solubility, toxicity, and synthetic accessibility, even before the primary high-throughput or virtual screening, which are primarily geared toward detecting potency alone. The best known physical property filters is Lipinski’s “rule-of-five”, which focuses on bioavailability [Lipinski-1997]. However, many other physical properties, such as solubility, atomic content, ring structures, and surface area ratios can also be considered. FILTER provides algorithms for calculating many of these properties and applying them with filters based on literature studies.

Finally, types of compounds that can be troublesome at later stages should be eliminated. For instance, Shoichet’s aggregating compounds often produce false positives that can waste enormous resources if they are identified by virtual or HTS ([McGovern-2003]; [Seidler-2003]). Similarly, dyes can appear to be inhibitors by interfering with colorimetric or fluorometric assays or by nonspecific binding to the target protein.

Variations of Filters

Different types of filters are appropriate under different circumstances. Very early in a project, when little or no information on the structure-activity relationship (SAR) is available, very strict drug-like filters can be applied. This avoids spending chemistry resources to pursue difficult compounds that may not be modifiable for appropriate properties. However, when considering compounds to purchase for HTS, different filters can be applied. Oprea et al. pointed out that the best molecules for initial HTS are smaller and have less functionality than drugs but show some activity [Oprea-2000]. Therefore, strict lead-like filters can be applied to ensure that hits identified from HTS can be investigated for potential leads, especially those that may be larger or more highly functionalized. However, when SAR suggests that particular compounds or series may yield valuable information, filtering criteria can be loosened, because secondary screens can be effective in detecting useful compounds. Circling back to the medical analogy, this is a case where an improved primary screen with a dramatically improved false-positive rate (e.g., 1 in 100,000) can be applied to a larger population without detrimental effects on the positive-predictive value.

FILTER provides the following filters:

  • BlockBuster: The BlockBuster filter is based on 141 best-selling, non-antibiotic prescription drugs. We designed the physical property portion of the filter so that it passes all of the compounds. The physical property values in this filter are quite good. However, the functional group filters in this filter may be too restrictive because 141 compounds will not span all acceptable functionality.

  • Drug: The original drug filter is provided as well [Oprea-2000], but has been shown to be very restrictive. The BlockBuster filter was developed as an alternative.

  • Fragment: The fragment filter is designed specifically to filter molecules with attachment points. This filter is used inside CHOMP to restrict the fragments generated for the BROOD database.

  • Lead: The Lead filter corresponds to the Oprea lead-like filters that are useful for preparing HTS screening databases.

  • PAINS: The PAINS filter is based on work by Baell which describes identification and filtering of promiscuous (nonspecific) actives across a number of screening types and targets [Baell-2010]. Unlike the other filter types, the PAINS filter consists of only functional group filters; no physical property filters are included. In practice, combination of the PAINS filter with a user-defined set of physical property filters may be desirable.

Hint

We recommend the default BlockBuster filter for most purposes. If you have specific filtering needs or are unsatisfied with the results, please see the filter file.

If you modify the filter file, the depictions found in the Functional Group Rules section can be particularly helpful in determining which functional groups are indicated by each name in the file.

Accumulation of Rules

FILTER contains rules to judge the quality of molecules on various properties. Individually, each of these rules seems quite reasonable and even profitable. However, when each molecule is tested against hundreds of individual filters, the fraction of molecules that pass all the filters can be surprisingly small. Sometimes less than 50% of vendor databases pass the filters. If this is unacceptable, we recommend you examine the predicted aggregator, solubility, and Veber filters. In our experience, these are the most common failures. The best method of investigating failures is looking at the filter log.

The BlockBuster filter addresses these challenges. Sometimes the range needs to be expanded in order to allow more compounds to pass. For each value, its properties extend from the 2.5th percentile to the 97.5th percentile. The differences between the original BlockBuster filter and the adjusted filter are both in reasonable ranges. For instance, the full range of molecular weights for the BlockBuster filter spans 130 to 781, while the 2.5th percentile is 145 and the 97.5th percentile is 570. Surprisingly, when the reduced filter is used, only 75 of the 141 original molecules pass the filter! This demonstrates how slight changes to many filters can lead to a significant reduction in the number of compounds that pass all filters.

The opposite approach of allowing everything to pass can be equally challenging. For example, a filter was designed around the “small molecule drug” file available from the DrugBank website. In order to pass all of these molecules, including some that would no longer be acceptable, many of the individual filters must be set to unreasonable values, such as a molecular weight range from 30 to 1500 and a heteroatom count range from 1 to 60.

Hint

Because every project is different, we recommend that you inspect each filter file. The depictions found in the Functional Group Rules section can be particularly helpful in discerning what functional groups are desirable and undesirable.

Filter Preprocessing

Before the application of any molecular property filters, a preprocessing step may be used to alter the molecule significantly to fit the criteria needed for most modeling applications. This filter preprocessing step is applied to each molecule in a series of stages in the following order:

  1. Metal Removal

  2. Salt Removal

  3. Canonicalization

  4. pKa Normalization

  5. Normalization

  6. Reagent Selection

  7. Type Checking

  8. MMFF94 Atom Type Checking

Metal Removal

Metal removal is the first stage of element-based filtering. This stage removes specified metal complexes from the molecule. It will not reject a molecule with a metal complex. This allows the filter to treat atoms in the counterion portion of a molecule separately from the atoms in the primary portion of the molecular record.

For instance, organic molecules that are complexed with silver can be eliminated based on their metal chelate even though they themselves are acceptable, while at the same time eliminating a sulfate counterion from another molecule before it leads to elimination of the acceptable cationic molecule.

Salt Removal

This step deletes all atoms that are not part of the largest connected component of a compound. This effectively eliminates all non-covalently bound portions of the compound.

Canonicalization

This step canonicalizes the atom and bond order of the portion of the molecule that remains after the previous removal steps. This is necessary to avoid different atom orderings producing slightly different normalizations in the following normalization steps.

pKa Normalization

pKa normalization uses a rule-based system to set the ionization state of input molecules. If pKa normalization is turned on, the molecule is set to its most energetically favorable ionization state for pH=7.4. The rule-based nature of this calculation allows it to be very fast. Furthermore, despite being rule-based, this approach takes into account many secondary charge interactions.

While more advanced levels of theory can be found for predicting ionization states, this method is very well suited to virtual-screening database preparation. However, this may not be appropriate for hit-to-lead or lead optimization.

Normalization

In addition to pKa normalization, FILTER allows any number of additional molecular normalizations. Since normalizations are usually specific to a particular company or site, FILTER provides the ability for users to input normalizations, such as the nitro tautomer state, but does not provide default implementations.

Reagent Selection

Reagent selection for small linear library synthesis or large combinatorial library synthesis is usually a necessary task. A user hoping to identify a set of acyl-halide reagents could specify a selection parameter to require each compound to have exactly one acyl-halide. In addition, the filter could be modified to exclude functional groups (such as primary amines) that may be acceptable for typical lead-like molecules, but are not acceptable for the specific reagent the user has in mind.

Therefore, the selection parameter is the reverse of a filtering parameter. The molecule must include the given substructure in order to pass the filter.

See also

The select parameter in filter files.

Type Checking

This checks the valence state and formal charge of the entire molecule. The check identifies molecules that are poorly specified, or represent nonsensical chemical states, often from corrupt input data. For example, an oxygen with eight hydrogens attached or a carbon with a +9 formal charge would be rejected.

MMFF94 Atom Type Checking

This checks that all atoms in the molecule have valid MMFF94 atom type assignments. The check identifies molecules that will fail downstream processing that depends on MMFF94 atom types (e.g., Omega).

Molecular Properties and Predictors

FILTER provides a range of properties to predictors to be used as molecular filters. Molecular properties are distinct physical properties that can be measured, including the following:

There have been several attempts to develop fast-approximate QSAR models for bioavailability. The first of these was Lipinski’s work [Lipinski-1997], followed by work at Pharmacopia [Egan-2000], Abbott [Martin-2005], and GSK [Veber-2002]. The simplest and probably most trusted work includes studies of LogP and PSA used by Egan. Martin’s Abbott Bioavailability Score (ABS) appears to be a refinement of the first generation models and is designed specifically to categorize a molecule’s probability of having a bioavailability >10% in rats.

See also

The Pharmacokinetic Predictors section in the Filter Files chapter.

Structural and Chemical Features

A number of important structural and chemical features of molecules should be limited for virtual screening. The following simple measures are provided as filterable properties:

  • Molecular weight

  • Ring count

  • Ring-system size

  • Size of non-ring structures

  • Length of unbranched chains

  • Heteroatom fraction

  • Halide fraction

  • Formal charges

  • Rotatable bonds

MolProp TK also includes slightly more complex algorithms for hydrogen bond donors and acceptors as well as chiral centers.

See also

The Basic Properties section in the Filter Files chapter.

Functional Groups

Functional group removal remains at the heart of the FILTER algorithm. Functional groups fall into several categories, including:

  • Reactive or labile groups

  • Undesirable groups

  • Generally acceptable groups

  • Protecting groups

  • User-derived groups

Reasonable defaults are provided for each functional group in the above categories, with the filter files being designed as basic guides. They can be tailored for each user’s particular needs. Examining the filter file can provide useful information. Developing a custom filter for a specific task is strongly encouraged.

See also

The filter files contain only functional group names. For clarification of the definitions, please refer to the Functional Group Rules section for complete descriptions.

Dyes

In the past, colored molecules interfered with many high-throughput assays. While assay technology continues to advance and this problem has decreased, most dyes are not commonly carried forward in lead development projects. Therefore, a pattern-based filter for dye molecules is included. While these patterns occasionally identify molecules that are considered acceptable by some users, in general, they identify molecules that the majority of chemists would rather not see at the top of their virtual screening hit lists.

LogP

The XLOGP algorithm ([Wang-R-1997]) is provided because its atom type contribution allows calculation of the XLogP contribution of any fragment in a molecule and allows minimal corrections in a simple additive form to calculate the LogP of any molecule made from combinations of fragments. Further, although the method contains many free parameters, its simple linear form allows for ready interpretation of the model, and most of the parameters in the model make rational sense.

Unfortunately, the original algorithm is difficult to implement as published. First, the internal hydrogen bond term was calculated using a single 3D conformation, which was found to be both arbitrary and unnecessary. This arbitrary 3D calculation has been replaced with a 2D approach to recognize common internal hydrogen bonds. In tests, the 2D method worked comparably to the published 3D algorithm. Next, the training set had a few subtle atom-type inconsistencies.

Both of these problems were corrected and refit to the original XLOGP training data. This implementation gives results that are quite similar to the original XLOGP algorithm, so it is called OEXLogP to distinguish it from the original method.

See also

LogS

The work of Samuel Yalkowsky [Yalkowsky-1980] at the University of Arizona has resulted in what is now called the “generalized solvation equation.” It states that the solubility of a compound can be broken into two steps, first the melting from the pure solid to pure liquid, and second, the phase transfer from pure liquid into water. For many small organic molecules, this second step is somewhat related to LogP.

Because of this relationship, we explored the use of the XLogP atom types in solubility prediction, with the goal of providing an approximate though robust and fast method for calculating solubility. We fit the XLogP atom types to a training set of nearly 1000 public solubility values, then derived a linear model for solubility. The model is extremely fast and can classify compounds as insoluble, poorly soluble, slightly soluble, moderately, soluble, or very soluble. However, it has difficulty in predicting the solubility of compounds with ionizable groups. Furthermore, it is not suitable for the PK predictions that come late in a project. Nevertheless, it can be utilized to eliminate compounds with severe solubility problems early in the virtual screening process.

See also

The solubility parameter in the Filter Files chapter.

Polar Surface Area

Topological polar surface area (TPSA) is based on the algorithm developed by Ertl et al [Ertl-2000]. In Ertl’s publication, use of TPSA both with and without accounting for phosphorus and sulfur surface area is reported. However, evidence shows that in most PK applications, it is better to avoid counting the contributions of phosphorus and sulfur atoms toward the total TPSA for a molecule. This implementation of TPSA allows either inclusion or exclusion of phosphorus and sulfur surface area, with the default being to exclude it.

See also

Including phosphorus and sulfur surface area:

Warning

TPSA values are mildly sensitive to the protonation state of a molecule. If the pKaNorm parameter is false, the TPSA value is calculated using the input structure, and if pKaNorm is true, the TPSA is calculated using the pKa normalized molecular structure.

See also

Lipinski and Hydrogen Bonds

The work of Lipinski [Lipinski-1997] introduced the application of simple filter-like rules to roughly predict late-stage PK properties, in particular oral bioavailability. Unfortunately, Lipinski’s rule of five has come into the common vernacular to such a large degree that some of the specific details are often lost. There are two critical examples of this. First, Lipinski used “violation of two rules” to categorize compounds. In the subsequent analysis, a significant difference in the two populations of molecules was detected, however, little analysis of the importance of a single violation was done. However, many now consider a single violation to be bad and two violations to be worse. Second, Lipinski used well-codified and well-understood yet imprecise definitions of “hydrogen bond donors” and “hydrogen bond acceptors” in his classification model. While this makes the algorithm quite understandable and easy to implement, it sometimes causes confusion with those who prefer more refined definitions of hydrogen bond donors and acceptors.

To address the first problem, users are allowed to set the number of Lipinski failures required to reject a molecule. In keeping with the original publication, the default value is 2. To address the second problem, two kinds of hydrogen bond donors and acceptors are calculated. In the Lipinski calculation, the published definitions are used (donor count is the number of nitrogen or oxygen atoms with at least one hydrogen attached and acceptors are the number of nitrogens and oxygens). For calculation of the number of donors and acceptors in the molecule for the sake of chemical properties, a more complex algorithmic approach is taken. This approach identifies the donors and acceptors outlined in the work of Mills and Dean [Mills-Dean-1996] and also in the book by Jeffrey [Jeffrey-1997].

See also

The Lipinski violations and hydrogen-bond acceptors sections in the Filter Files chapter.

Aggregators

The Shoichet lab ([McGovern-2003]; [Seidler-2003]) has demonstrated the importance of small-molecule aggregation in medium and high-throughput assays. Because these small-molecule aggregates can sequester some proteins, they give the appearance of being active inhibitors. There are now several hundred published aggregators, in addition to a published QSAR model for predicting aggregation propensity. The user can eliminate any of the known aggregators and also eliminate compounds that are predicted to be aggregators using the QSAR model. The published QSAR model for predicting aggregators may be considered quite aggressive, as it occasionally identifies compounds that are known to be genuine small-molecule inhibitors of a specific protein. While in most cases this model can be useful, it is important to gain experience with the model and judge the performance for specific needs.

In some cases, aggregation properties are specific to specific experimental conditions, thus caution should be exercised in the interpretation of these predictions. Nevertheless, aggregation remains an important issue in HTS hit follow-up and, short of experimental validation, this flag may be the best available.

See also

The aggregators parameter in the Filter Files chapter.

PAINS

Pan-assay interference compounds (PAINS) filters are a set of functional group filters created to identify and filter promiscuous (nonspecific) actives across a number of screening types and targets [Baell-2010]. The original PAINS paper includes a set of functional group prefilters, created via two iterations of their work, and three sets of structure class filters, called filter sets A, B, and C, which were used to identify and remove promiscuous actives. The OpenEye PAINS filter set is a combination of all four filter sets (the functional group prefilter set plus the three structural class sets) into a single filter set.

There are two distinct use cases for the PAINS filters. The first is the typical prefiltering of structures prior to screening or compound selection to remove undesirable molecules. One can simply eliminate any molecules which fail the PAINS filtering.

The second use case, which may ultimately be more interesting, is to use the PAINS filters retrospectively on a set of data which has been filtered (and tested) using another methodology. In this scenario, the PAINS filters can be used to identify potentially problematic reactive functionality in otherwise reasonable, drug-like structures. By using the table output mode and identifying matches for the A, B, and C filter sets, one can identify fairly specific chemical functionalities that have been known to result in nonspecific activity. This can be useful information for either prioritizing further lead development or for modifying lead structures to mitigate adverse reactivity.

There are approximately 650 SMARTS rules which have been converted from the SLN query strings provided in the supplemental information of the original paper. For the functional group filters, the authors do not provide identifying names for each individual filter, so our PAINS rules are named based on the comments used to describe each functional group class, prefixed with the string pains_fg_[classname]. For the structural class filters, the authors provided unique RegID values for each filter. For those cases, our rules are named based on the PAINS group and RegID as pains_[group]_[RegID].

Note that in certain cases, the SLN queries cannot be represented with a single SMARTS, so in those cases the rules have been split out into multiple SMARTS patterns and are named pains_[group]_[RegID]1, pains_[group]_[RegID]2, etc.

Hint

Unlike the other filter types, the PAINS filter consists only of functional group filters; no physical property filters are included. It may be desirable to combine the PAINS filter with a separate user-defined set of physical property filters. The filter file section describes the creation of additional filter rules.