PharmCast SCP v10
What a molecule can present, not what it looks like
A pharmacophore fingerprint records the three-dimensional arrangement of binding features a molecule can present, so two compounds from completely different scaffolds can be compared on the thing a protein actually reads. That encoding is the 10,549-bit three-point PharmPrint™, and comparing two of them is PharmSim™. PharmCast is the model family that predicts a PharmPrint straight from a two-dimensional structure, so a full comparison of two molecules runs without ever building a conformer, and agrees with the reference calculation at Pearson 0.980.
This page reports PharmCast SCP v10, the current model. SCP names the three sources it is trained on, Screening collection, ChEMBL and Peptides: 5,887,229 molecules used for training, from a frozen snapshot of 5,946,696. It supersedes the earlier collection-only and collection-plus-peptide models.
The fingerprint
Every pharmacophore in the scheme is a triangle of three features at three measured distances. Enumerate every combination of feature types and distance ranges and you get a fixed vocabulary of possible triangles. Each one is a bit. A molecule sets the bit for every triangle it can actually form.
2.0 to 4.5, 4.5 to 7.0,
7.0 to 10.0, 10.0 to 14.0, 14.0 to 19.0 and
19.0 to 24.0 Å.Because the vocabulary is geometric rather than substructural, two molecules with nothing in common on paper can score highly against each other if they present their features in the same places. That is exactly the comparison a scaffold hop needs, and it is not what a two-dimensional fingerprint measures.
PharmPrint and PolyPharmPrint
A molecule is not one shape. The original method, published by McGregor and Muskal in 1999 and 2000, built many conformers, fingerprinted each, and combined them into a single ensemble fingerprint by taking the union of all the bits. That answers whether a molecule can present a pharmacophore somewhere in its accessible landscape. What it cannot tell you is which shape did it.
Original PharmPrint
Build many conformers, fingerprint each, then OR all the bits into one combined fingerprint.
The combined fingerprint answers what the molecule can present. The conformer that produced any given match is lost, so an alignment has to search the conformational space again from scratch.
PolyPharmPrint
Build hundreds of conformers and keep every fingerprint separate, so a match is between a specific conformer of the query and a specific conformer of the hit.
The matching geometry is recoverable. Alignment starts from the conformer pair that actually scored, which turns an expensive global search into a targeted one.
Keeping the conformers separate is what makes the fingerprint useful downstream rather than only as a filter. It is the same idea that lets the ChIP me too campaign show a designed molecule overlaid on its reference in the exact pose that produced the score.
Turbocharging the comparison
A full pharmacophoric comparison of two molecules, starting from nothing but their SMILES, runs without building a three-dimensional conformer at all. The answer agrees with the reference calculation at Pearson 0.980, and in a retrieval test the true nearest neighbor lands in the surrogate's top 100 for 89.3% of queries, and its top 500 for 97.5%. Measured over 22,793 queries against 22,793 candidates the model has never seen.

How it works
Two real molecules the model has never seen, carried through every stage to the number you actually use. The fingerprint is the intermediate; the pharmacophoric similarity is the product.

CC(C)(C)NC(=O)[C@@H]1C[C@@H]2CCCC[C@@H]2CN1C[C@@H](O)[C@H](Cc1ccccc1)NC(=O)[C@H](CC(N)=O)NC(=O)c1ccc2ccccc2n1indinavirCC(C)(C)NC(=O)[C@@H]1CN(Cc2cccnc2)CCN1C[C@@H](O)C[C@@H](Cc1ccccc1)C(=O)N[C@H]1c2ccccc2C[C@H]1OWhat it is for
PolyPharmPrint™ captures the pharmacophore at the level of the individual conformer, where the historic PharmPrint combined a molecule’s conformers into a single ensemble fingerprint for rapid screening. What is new here is a rapid screening route of a different kind: a neural embedding that returns the fingerprint, and with it the comparison, without building the conformers at all. Almost everything useful you do with one is comparative: is this design close to that reference, which hundred of these million compounds are worth buying, which of these hits are really the same chemotype. Those are ranking and clustering questions, and they depend on the similarity between two fingerprints rather than on any single fingerprint being perfect.
The conventional route to that number is expensive and has two variants, both of which start the same way. You generate a three-dimensional conformer ensemble, then fingerprint it. Either you OR the conformers together into one ensemble fingerprint, or you keep them separate and carry a fingerprint per conformer, as the polypharmprint workflow does so the best matching conformer pair can be found. Either way you pay for the conformers, and then you pay again for fingerprinting each one.
This model bypasses all of it. No conformer generation, no per-conformer fingerprinting, no ensemble to assemble. It reads a canonical SMILES and predicts the fingerprint directly, one output for every one of the 10,549 bits. To be precise about what it predicts: it was trained on ensemble fingerprints, the OR of 100 conformers per molecule, so what it reproduces is the ensemble fingerprint and the similarity between two of them, and the results below measure it on the operation that matters: does its answer to "how similar are these two" match what the real calculation says, and does it rank candidates in the same order.
It is a filter, not a replacement. The real calculation remains the arbiter for anything you act on. The surrogate is what lets you put far more chemistry in front of it.
| Input | canonical SMILES |
| Features | 2,048-bit Morgan radius 2, plus 11 descriptors |
| Output | all 10,549 bits, one neuron each |
| Trained on | 5,887,229 molecules: the screening collection, ChEMBL and protein loop peptides |
| Tested on | 155,648 catalog and 139,700 ChEMBL molecules held out of the training set, plus the 13,500 molecule reserved peptide test set |
| Median agreement per molecule | 0.881 on screening collection chemistry, 0.914 on the peptide test set |
| Similarity correlation | 0.980 on screening collection chemistry |
| Screening rate | 2,644 compounds per second, six worker processes |
Where the data comes from
The screening collection. The June 2026 Enamine screening collection, physically stocked material rather than a virtual enumeration. Property filters on ingest kept 4,612,044 of the 4,774,670 compounds. It is broad in frameworks and dense in analogs at the same time: 135,768 distinct Bemis-Murcko scaffolds in a 200,000 molecule sample, 88% of them appearing once, against a median nearest-neighbor Tanimoto of 0.714 with 64,473 compounds carrying a 2D-identical twin. Every fingerprint here is built the same way from it: 100 ETKDGv3 conformers per compound, UFF optimized, each fingerprinted and ORed into one record per molecule.
The protein loop set. SCP v10 also trains on peptides taken from experimentally determined protein structures, 1,524,535 loop instances read from 69,450 usable entries, capped with their real flanking atoms and ensemble enhanced with every crystallographic conformation observed for each sequence. Those collapse to the 136,494 distinct loops of two to six residues the fingerprints are built from, all from structures at 2.5 Å or better and R-free 0.23 or better.
Full profile of both source sets →
The test set
Trained on 5,887,229 molecules across the three source sets, and evaluated on three populations none of which it has seen: 155,648 real, purchasable June 2026 Enamine compounds the training-set ingest filter excluded, checked by canonical SMILES against all 4,617,292 screening collection structures; 139,700 activity-backed ChEMBL molecules held out of the version 10 training set; and the 13,500 loop peptides reserved for testing. That is 308,848 molecules scored in total.
What the model is entitled to be asked
A critical assumption underlies every number on this page. Every molecule used to train and to evaluate this model is one that has actually been made or actually observed: compounds held in a commercial screening collection, and peptides extracted from experimentally determined protein structures. The model has only ever been shown chemistry that is synthetically real, so a degree of synthetic feasibility is built into its training distribution rather than imposed on it afterwards.
An arbitrary 2D structure that has never been synthesized is therefore outside the domain the model was fitted on. It is not necessarily excluded, and the model will return an answer for it, but that answer is an extrapolation and should be treated as one.
One test set, three chemistries
Catalog chemistry is not the only thing the model is asked to handle, so
the held out set above is widened here to everything it meets in practice:
molecules from the screening collection fingerprinted after the training
snapshot, loop peptides whose names the model never saw, and real ChEMBL compounds above 600 molecular weight, each carrying a measured
activity value, which lie outside the size range the collection covers. Scored by the
composite model, pharmcast_scp_v10.pt, trained on
4,609,488 collection molecules, 1,214,214
ChEMBL compounds and 122,994 peptides. Every point is a pair.

| Population | Median error | Within 0.05 | Pearson | Ranking accuracy |
|---|---|---|---|---|
| loop peptides, reserved test set | 0.016 | 88% | 0.984 | 0.952 |
| ChEMBL, activity backed | 0.027 | 75% | 0.936 | 0.889 |
| catalog chemistry | 0.008 | 89% | 0.980 | 0.936 |
| reference calculation against itself | 0.006 | 0.995 | ceiling |
Catalog chemistry lands on the diagonal at 0.008 median error and a correlation of 0.980, within about a third of the reference calculation's own reproducibility of 0.006. ChEMBL is looser, at 0.027, with roughly one pair in four outside 0.05.
The peptide row is the reserved peptide test set: 13,500 loops from PDB entries at 2.5 Å or better and R-free 0.23 or better, excluding every entry used to build the training set and every loop whose amino acid sequence already appears in it. Those peptides are novel by sequence rather than merely by canonical SMILES, so a poor score there could not have been blamed on soft coordinates. Median per molecule agreement on that set is 0.914, against 0.881 on screening collection chemistry and 0.860 on ChEMBL.
The peptide result is what training on peptides bought: this model saw 122,994 of them. ChEMBL is the weakest population, and on the 573 real ChEMBL compounds above 600 Da that were held out of every model, ranking accuracy is 0.826 against 0.934 on screening collection chemistry. The cloud there is centered on the diagonal rather than displaced from it, so the model is not biased on those molecules so much as imprecise about them. Read the combined figure as a statement about coverage rather than as a single accuracy: the 0.016 median error over all three is a blend of two chemistries it handles well and one it handles more coarsely.
Where it holds up, and where the accuracy comes from
Coloring the same pairs by how alike the two molecules are in two dimensions answers a question the headline number cannot.

It asks whether the surrogate is simply re-deriving 2D similarity. It is not. A pair of molecules that look nothing alike on paper is predicted about as well as a pair that look somewhat alike. Under SCP v10 the correlation within each 2D band runs 0.966, 0.981, 0.973, 0.959 and 0.942, from the least two-dimensionally similar pairs to the most, over 25,165 pairs with 6,000 in each of the four lower bands and 1,165 above 0.50. That is the property that matters for scaffold hopping.
Mass is where coverage runs out rather than where the method does. The screening collection is filtered at 600 Da on ingest, so everything the model has seen above that line comes from ChEMBL and from the peptides: 91,388 molecules, 1.54% of the training set. That is why error is highest on ChEMBL and why it rises with mass.
How closely each single fingerprint matches
Everything above is about pairs. This is about one molecule at a time. A pharmacophore fingerprint is 10,549 bits, each one either present or absent, so the surrogate's output for a single molecule is 10,549 true or false predictions against a known answer. The Matthews correlation coefficient scores exactly that. It uses all four cells of the confusion matrix, so it cannot be flattered by the fact that most bits are off.
One is a perfect match and zero is chance.

All three medians are measured on molecules the model never trained on: 155,648 catalog, 139,700 ChEMBL and 13,500 loop peptides from the reserved test set. The catalog and ChEMBL populations are held out by time, every molecule fingerprinted after the seal; the peptides are a sequence-novel test set permanently excluded from training.
Each population is single-peaked and tight, so the surrogate behaves consistently within a chemistry rather than splitting it into molecules it handles and molecules it does not. The populations separate from each other, which is what the numbers say: catalog chemistry is what the training set is mostly made of, and agreement falls as a molecule moves away from it.
Searching new chemistry
The usual job is not to reproduce a calculation you could have run. It is to score designs nobody has evaluated, far more of them than you can afford the real calculation on, and spend the next round's effort on the promising ones.
What is being enriched here is pharmacophoric similarity to the query, not bioactivity. Nothing in this measurement knows anything about a target, an assay or a hit. "The closest 1%" means the 1% of the collection with the highest real pharmacophoric similarity to the query molecule, as the full conformer calculation scores it.
So: rank a held-out pool by the surrogate, walk down that ranking, and record how much of that genuinely closest chemistry you have picked up at each depth.

Put as a retrieval question, over 22,793 candidates and 22,793 queries: the true nearest neighbor has a median rank of 5 under the surrogate, and it falls inside the surrogate's top 100 for 89.3% of queries and its top 500 for 97.5%. That is the operationally relevant behavior, because the surrogate is used to build a shortlist that the reference calculation then adjudicates.
What this saves you
Getting a real pharmacophore fingerprint from a SMILES means generating a conformer ensemble and then fingerprinting it. Almost all of the cost is the ensemble: the bit calculation itself is a small fraction of it, and that cost is unavoidable in the conventional route whether you OR the conformers into one fingerprint or keep them separate for polypharmprint-style best-pair matching. Both variants pay it.
The surrogate never generates a conformer and never fingerprints one. That single omission is where the speed comes from, and it is a change of kind rather than a change of degree.
That change is what makes exhaustive exploration possible. When a fingerprint costs a conformer ensemble you budget your chemistry: a genetic algorithm scores a few thousand designs a generation, and screening a large collection is a project rather than a query. When it does not, you stop budgeting. A CHIP-style search can enumerate and score millions of designs per generation instead of thousands, and an entire commercial collection becomes something you scan in a minute and rescan whenever the question changes.
Concretely, that opens up:
- Design search. A CHIP-style genetic algorithm can enumerate and score millions of candidate designs per generation instead of thousands, so the search covers chemistry it previously could not afford to look at.
- Screening collection triage. Rank an entire commercial catalog against a reference in 29.3 minutes, where the conventional route needs 3,697 core hours, then send only the top slice to the reference calculation.
- Clustering and diversity. All-pairs pharmacophoric comparison over a large set becomes tractable, which makes 3D-aware clustering and diversity selection possible at collection scale.
- Scaffold hopping. Because accuracy does not depend on 2D similarity, the ranking surfaces molecules that a 2D fingerprint would never retrieve.
The real calculation stays the arbiter. The surrogate is what lets you put far more chemistry in front of it.
How the timings were measured. All timings on this page were measured on one machine, an Apple M3 Ultra with 28 CPU cores, a 60-core GPU and 256 GB of unified memory, running macOS 26.5.1 with PyTorch 2.9.1 on CPU. The GPU was not used, and both routes were timed on the same hardware, so the comparison is like for like. Per molecule cost depends on the batch regime: 0.288 ms in a library-scale batch against 0.357 ms processed alone.
Where the per conformer calculation still earns its place
The surrogate predicts an ensemble fingerprint, which answers which molecules are pharmacophorically close. It cannot tell you which conformation of them does the work, because the ensemble has already been collapsed. That question still needs the real per conformer calculation, and the two fit together rather than competing.
| Stage | Method | Scale | Answers |
|---|---|---|---|
| 1. Find | Turbocharged pharmacophoric similarity | millions | Which molecules resemble the query in pharmacophore space |
| 2. Deconvolute | Per conformer fingerprints, best matching pair | the survivors | Which conformation of each carries the overlap, and what it looks like |
Stage one scans an entire collection against a query in 29.3 minutes and hands back a ranked shortlist. Stage two takes only that shortlist, generates real conformers, fingerprints each one separately rather than ORing them, and scores every conformer of the query against every conformer of the hit. That surfaces the specific conformer pair carrying the similarity, which is the pose you can render, inspect and take into design.
Doing stage two across the whole collection would be prohibitive. Doing it across a few hundred survivors is routine. The surrogate is what makes the shortlist small enough and good enough to be worth the expensive look.
Why this matters for design
The cost of scoring is what bounds how much chemical space a pharmacophore-driven search can afford to cover. Our ChIP de novo design engine scores candidate molecules on pharmacophore similarity to a reference, so making that score roughly four orders of magnitude cheaper changes how much route and building-block space a campaign can explore before it has to commit. The worked me too campaign is one example of what the score is used for.
Reverse peptide mimetics → PharmCast fingerprints on both sides of the comparison: every co-crystal ligand with a measured potency scored against all 3,368,420 capped peptides of one to five residues. The forty closest pairs reach a pharmacophore Tanimoto of 0.916 with almost no chemical similarity, each with a rotatable superposition.Talk to us about pharmacophore search
PolyPharmPrint runs against your compounds and your references, and the surrogate makes it practical at collection scale. We are happy to walk through the method, the validation, and where the limits sit.
Muskal, S. M.; McGregor, M. J. PharmCast: rapid generation of three-dimensional pharmacophore fingerprints from two-dimensional structure without conformer generation. bioRxiv 2026, doi: 10.64898/2026.09.02.748999. · McGregor, M. J.; Muskal, S. M. Pharmacophore fingerprinting. 1. Application to QSAR and focused library design. J. Chem. Inf. Comput. Sci. 1999, 39, 569–574. · McGregor, M. J.; Muskal, S. M. Pharmacophore fingerprinting. 2. Application to primary library design. J. Chem. Inf. Comput. Sci. 2000, 40, 117–125. Surrogate figures and figures quoted on this page are computed from the evaluation artifacts of the trained model. The method is described here in full; the trained model itself is not distributed.