Eidogen-Sertanty

PharmCast

What a molecule can present, not what it looks like, predicted straight from the 2D structure, with no conformers at all.

A pharmacophore fingerprint records the three-dimensional arrangement of binding features a molecule can present, so two compounds from completely different scaffolds can be compared on the thing a protein actually reads. That encoding is the 10,560-bit three-point PharmPrint, and comparing two of them is PharmSim.

Its cost has never been the fingerprint. Of the 2.48 seconds needed to fingerprint one catalogue molecule over 100 conformers, 2.44 s is conformer generation and only 0.043 s is the fingerprint itself. PharmCast is a model that predicts the complete fingerprint directly from a SMILES string, removing the conformational stage rather than accelerating it. That is why the speed-up is of a different order to what optimisation normally buys.

0.034 msOne complete comparison of two molecules, from structure alone, against 5.0 s conventionally
0.97Correlation with the real calculation on catalogue chemistry
0.895Median per-molecule bitwise agreement (MCC) on 9,975 held-out molecules
2.6 MMolecules in the training corpus, every one fingerprinted with the real calculation

The fingerprint

A fixed geometric vocabulary, not a substructural one

Every pharmacophore in the scheme is a triangle of three typed features at three measured distances. Enumerate every combination of feature types and distance ranges and you get a fixed vocabulary of possible triangles. Each one is a bit. A molecule sets the bit for every triangle it can actually form.

A three-point pharmacophore: three typed features p1, p2, p3 joined by three measured distances d1, d2, d3, beside the seven feature types and six distance bins that together enumerate 10,560 bits.
One pharmacophore. Seven feature types (acceptor, donor, negative, positive, hydrophobic, aromatic and other) at six binned distances from 2.0 to 24.0 Å. Every legal triangle is one bit of the fingerprint.

Because the vocabulary is geometric rather than substructural, two molecules with nothing in common on paper can score highly against each other if they present their features in the same places. That is exactly the comparison a scaffold hop needs, and it is not what a two-dimensional fingerprint measures.

Two different molecules, each shown in three dimensions with the same acceptor, donor and ring features picked out and the three inter-feature distances measured between them.
The same typed triangle found in two unrelated structures. The distances, not the atoms, are what the fingerprint records.

What PharmCast predicts

PharmCast reads a canonical SMILES and predicts all 10,560 bits, one output per bit. It was trained on ensemble fingerprints, the bitwise OR of 100 conformers per molecule, so what it reproduces is the ensemble fingerprint and the similarity between two of them. The output is the standard native format: the same 330 unsigned 32-bit words the real calculation emits, in the same bit convention, so existing PharmSim tooling consumes it unchanged.

Input

Structure
Canonical SMILES
Features
2,048-bit Morgan radius 2, plus 11 descriptors

Output

Fingerprint
330 × uint32: 10,560 bits, 10,549 real pharmacophores
Format
Native .pfp, byte-identical convention
Two rules for using the scores

Never predict the reference of a campaign. Fingerprint it once with the real calculation, because there is no reason to accept model error on the one molecule the work is aimed at. And PharmCast scores rank candidates; they are not quoted as the similarity. The predicted fingerprint carries a small systematic offset in bit count, harmless for ordering and misleading in a report. Published numbers come from a real rescore of the survivors.

Accuracy

Measured on molecules the model has never seen

Four scatter panels of PharmCast similarity against real pfpall similarity: screening collection, loop peptides, large ChEMBL compounds, and all three together. The first two hug the diagonal tightly; the large-compound panel is visibly more scattered.
PharmCast similarity against the real calculation on three distinct chemistries, 1,200 pairs each. The first two panels sit on the diagonal. The third, real ChEMBL compounds above 600 Da, plainly does not, and the reason is in the next section.
RegimeMedian errorCorrelation rPairwise ranking
Catalogue chemistry0.020.9792%
Loop peptides0.020.97n/a
Large compounds, above 600 Da0.070.6371%
The real calculation against itself0.0060.995ceiling

That last row is the one that sets the scale. Rebuilding the same molecules with a different embedding seed reproduces pair similarity to 0.006 at r 0.995, so the reference calculation agrees with itself an order of magnitude more tightly than PharmCast agrees with it. The ground truth is not the problem, and 0.006 is the floor no surrogate can beat.

Hexbin density plot of surrogate pharmacophore similarity against real pharmacophore similarity over 30,000 molecule pairs, tightly following the diagonal, Pearson 0.96 and Spearman 0.96.
30,000 pairs of catalogue molecules. Each point is one pair: the similarity the surrogate reports against the similarity the real calculation reports.

Where it holds, and where it does not

Applicability domain

PharmCast SP is calibrated for catalogue-like chemistry up to about 600 Da, plus peptides to about 900 Da. Outside that range it is extrapolating, and the error profile above shows exactly what that costs.

The reason is entirely in the training corpus, and it is worth stating plainly. The screening collection is filtered at 600 Da on ingest, so it contributes nothing above that line by construction. Only about 37,060 training molecules, 1.43% of the corpus, sit above 600 Da, and every one of them is a peptide. Above 800 Da it is 0.10%. So when PharmCast is asked about a large non-peptide drug it is extrapolating from a few thousand peptides, and its error rises from 0.02 to 0.07 while r falls from 0.97 to 0.63. This is a genuine size effect and not simply a harder task: a matched-spread catalogue control gives 92% on the identical protocol.

CorpusMolecules trained onShareMass range
Screening collection2,511,44096.69%142–598 Da, median 344
Loop peptides, ensemble enhanced86,0393.31%116–988 Da, median 576
Total2,597,479100%

What it is trained on

Every molecule fingerprinted with the real calculation, not a shortcut

The screening collection

The source is the June 2026 release of the Enamine screening collection, which currently states 4,774,670 compounds. These are real, orderable material: compounds synthesised and held in stock, quality controlled to at least 90% purity, not a virtual enumeration. Enamine's much larger combinatorial spaces are not included.

Property filters on ingest keep compounds up to 600 Da, at most 10 rotatable bonds and no more than one Lipinski violation; 4,612,044 survive. That filtered set is the index PharmCast draws from, and it is being fingerprinted continuously with the real 100-conformer calculation. SP v4 was trained on the 2,511,440 finished at the time of its snapshot. The 600 Da filter is applied here, on ingest, which is precisely why the model has no catalogue chemistry above that line to learn from.

Property5th pctMedian95th pct
Molecular weight255344465
cLogP0.72.74.7
Fraction sp3 carbon0.090.350.71

This is a lead-like library rather than a fragment or biologics-adjacent one. The median compound weighs 344, carries three rings of which two are aromatic, five rotatable bonds, and sits at cLogP 2.7. Only 8% carry defined stereochemistry, so the collection is predominantly flat, achiral scaffolding decorated with substituents.

Wide in frameworks, deep in analogs

Two measurements of the same collection pull in opposite directions, and both are true. Reduced to Bemis-Murcko scaffolds it is broad: 679 distinct frameworks per thousand molecules, 88% of them appearing exactly once, and the ten commonest together account for only 4.0%. But compared whole-molecule against whole-molecule (exhaustively, all 4,612,044 against each other with no sampling), it is dense.

0.714Median 2D similarity to the nearest other compound in the collection
54.9%Have a neighbour at 0.70 or closer
64,473Have a 2D-identical twin

The resolution is that the collection is wide in frameworks and deep in analogs around them: many scaffolds, each decorated many times over with closely related substituents. That has a direct consequence for how any model on this data must be evaluated: a random split places near-identical analogs on both sides and reports an optimistic number.

One caveat on that measurement

All of it is 2D structural similarity. Two compounds that look alike on paper can present different three-dimensional pharmacophores, and two that look unrelated can present similar ones. It describes the catalogue, not what the molecules do, which is the entire reason the 3D fingerprint exists.

The peptide corpus, and why it is not just more molecules

An RCSB query at 2.5 Å resolution or better with R-free at or below 0.22 returned 71,018 entries, of which 69,450 yielded usable fragments: 1,524,535 loop instances across 135,408 chains, collapsing to 117,741 distinct sequences and 133,136 distinct capped molecules. Loops are runs of two to six residues lying between annotated secondary structure, read from the helix and strand records in the mmCIF itself rather than re-derived.

Each fragment is capped with its real flanking atoms (an acetyl built from the preceding residue, an N-methylamide from the following one), so it never presents a free amine and acid the protein does not have. Backbone continuity is enforced by requiring the carbon-to-nitrogen distance between consecutive residues to fall between 1.0 and 2.0 Å rather than trusting residue numbering, which rejected 54,520 candidates that numbering alone would have accepted.

Ensemble enhancement is the novel step. Each distinct peptide gets one fingerprint that ORs together every rigid crystallographic conformation actually observed for it in the PDB with 100 computed conformers, generated exactly as for the screening collection. The conformer count therefore runs above 100, and the fingerprint carries both what nature was caught doing and what the molecule can do.

Repeats are kept rather than deduplicated, and the reason is measurable. Across 1,500 sequences seen in three or more entries, the median agreement between observations of the same sequence is 0.878, but 18% fall below 0.70, carrying genuinely different pharmacophores from one crystal structure to the next, and only 19% are near copies. Collapsing them to one representative would throw that away.

Seven deposited conformations of the loop sequence LGGK overlaid, beside a histogram of pharmacophoric similarity between pairs of those observed conformations, spread from 0.28 to 0.97 with a median of 0.56.
One loop sequence, seven deposited conformations, and the pharmacophoric similarity between them. A single sequence is not a single shape, which is the whole argument for the ensemble.

Model card

Release
PharmCast SP v4, trained 21 August 2026
Architecture
Feedforward network, 2,059 inputs to 10,560 sigmoid outputs
Training
40 epochs, 232.8 minutes, 2,597,479 molecules
Validation
9,975 held-out collection molecules, fixed across every release so versions stay comparable
Median bitwise MCC
0.8954
Applicability
Catalogue chemistry to ~600 Da; peptides to ~900 Da
Naming
S is the screening collection, P is peptides. There is no peptide-only model.

In progress not yet released

A corpus of real, activity-backed ChEMBL molecules in the 600–1000 Da band is being fingerprinted to close exactly the gap described above; 134,765 molecules carry a measured pChEMBL value and qualify. The model that includes them will be PharmCast SCP: Screening collection, ChEMBL, Peptides. No results are claimed for it here, and none should be assumed until it is measured on the same held-out set.

Using it

The Python package is small and self-contained. words_batch is the API that matters: the model earns its speed in a batch, and a single call is dominated by featurisation overhead.

from pharmcast import PharmCast, pharmtan, pharmsim

pc = PharmCast.load("pharmcast_sp_v4.pt")
a, b = pc.words_batch(["CCO", "c1ccccc1O"])   # 330 uint32 each

pharmtan(a, b)          # the PharmSim coefficient
pharmsim(a, b)          # ...with the shared and exclusive bit sets shown

Utilities for going from 2D structure to a native fingerprint, for PharmSim / PharmTan comparison, and for expanding a fingerprint into its 0/1 bit sets will be released as an open-source package. The reference fingerprint generator is not part of that release.

Citation

A preprint describing PharmCast is in preparation. In the meantime, the method it builds on is:

McGregor, M. J. and Muskal, S. M. Pharmacophore fingerprinting. 1. Application to QSAR and focused library design. J. Chem. Inf. Comput. Sci. 1999.
McGregor, M. J. and Muskal, S. M. Pharmacophore fingerprinting. 2. Application to primary library design. J. Chem. Inf. Comput. Sci. 2000.