A pharmacophore fingerprint records the three-dimensional arrangement of binding features a molecule can present, so two compounds from completely different scaffolds can be compared on the thing a protein actually reads. That encoding is the 10,560-bit three-point PharmPrint™, and comparing two of them is PharmSim™.
Its cost has never been the fingerprint. Of the 2.48 seconds needed to fingerprint one catalogue molecule over 100 conformers, 2.44 s is conformer generation and only 0.043 s is the fingerprint itself. PharmCast is a model that predicts the complete fingerprint directly from a SMILES string, removing the conformational stage rather than accelerating it. That is why the speed-up is of a different order to what optimisation normally buys.
The fingerprint
A fixed geometric vocabulary, not a substructural one
Every pharmacophore in the scheme is a triangle of three typed features at three measured distances. Enumerate every combination of feature types and distance ranges and you get a fixed vocabulary of possible triangles. Each one is a bit. A molecule sets the bit for every triangle it can actually form.
Because the vocabulary is geometric rather than substructural, two molecules with nothing in common on paper can score highly against each other if they present their features in the same places. That is exactly the comparison a scaffold hop needs, and it is not what a two-dimensional fingerprint measures.
What PharmCast predicts
PharmCast reads a canonical SMILES and predicts all 10,560 bits, one output per bit. It was trained on ensemble fingerprints, the bitwise OR of 100 conformers per molecule, so what it reproduces is the ensemble fingerprint and the similarity between two of them. The output is the standard native format: the same 330 unsigned 32-bit words the real calculation emits, in the same bit convention, so existing PharmSim tooling consumes it unchanged.
Input
- Structure
- Canonical SMILES
- Features
- 2,048-bit Morgan radius 2, plus 11 descriptors
Output
- Fingerprint
- 330 × uint32: 10,560 bits, 10,549 real pharmacophores
- Format
- Native
.pfp, byte-identical convention
Never predict the reference of a campaign. Fingerprint it once with the real calculation, because there is no reason to accept model error on the one molecule the work is aimed at. And PharmCast scores rank candidates; they are not quoted as the similarity. The predicted fingerprint carries a small systematic offset in bit count, harmless for ordering and misleading in a report. Published numbers come from a real rescore of the survivors.
Accuracy
Measured on molecules the model has never seen
| Regime | Median error | Correlation r | Pairwise ranking |
|---|---|---|---|
| Catalogue chemistry | 0.02 | 0.97 | 92% |
| Loop peptides | 0.02 | 0.97 | n/a |
| Large compounds, above 600 Da | 0.07 | 0.63 | 71% |
| The real calculation against itself | 0.006 | 0.995 | ceiling |
That last row is the one that sets the scale. Rebuilding the same molecules with a different embedding seed reproduces pair similarity to 0.006 at r 0.995, so the reference calculation agrees with itself an order of magnitude more tightly than PharmCast agrees with it. The ground truth is not the problem, and 0.006 is the floor no surrogate can beat.
Where it holds, and where it does not
PharmCast SP is calibrated for catalogue-like chemistry up to about 600 Da, plus peptides to about 900 Da. Outside that range it is extrapolating, and the error profile above shows exactly what that costs.
The reason is entirely in the training corpus, and it is worth stating plainly. The screening collection is filtered at 600 Da on ingest, so it contributes nothing above that line by construction. Only about 37,060 training molecules, 1.43% of the corpus, sit above 600 Da, and every one of them is a peptide. Above 800 Da it is 0.10%. So when PharmCast is asked about a large non-peptide drug it is extrapolating from a few thousand peptides, and its error rises from 0.02 to 0.07 while r falls from 0.97 to 0.63. This is a genuine size effect and not simply a harder task: a matched-spread catalogue control gives 92% on the identical protocol.
| Corpus | Molecules trained on | Share | Mass range |
|---|---|---|---|
| Screening collection | 2,511,440 | 96.69% | 142–598 Da, median 344 |
| Loop peptides, ensemble enhanced | 86,039 | 3.31% | 116–988 Da, median 576 |
| Total | 2,597,479 | 100% |
What it is trained on
Every molecule fingerprinted with the real calculation, not a shortcut
The screening collection
The source is the June 2026 release of the Enamine screening collection, which currently states 4,774,670 compounds. These are real, orderable material: compounds synthesised and held in stock, quality controlled to at least 90% purity, not a virtual enumeration. Enamine's much larger combinatorial spaces are not included.
Property filters on ingest keep compounds up to 600 Da, at most 10 rotatable bonds and no more than one Lipinski violation; 4,612,044 survive. That filtered set is the index PharmCast draws from, and it is being fingerprinted continuously with the real 100-conformer calculation. SP v4 was trained on the 2,511,440 finished at the time of its snapshot. The 600 Da filter is applied here, on ingest, which is precisely why the model has no catalogue chemistry above that line to learn from.
| Property | 5th pct | Median | 95th pct |
|---|---|---|---|
| Molecular weight | 255 | 344 | 465 |
| cLogP | 0.7 | 2.7 | 4.7 |
| Fraction sp3 carbon | 0.09 | 0.35 | 0.71 |
This is a lead-like library rather than a fragment or biologics-adjacent one. The median compound weighs 344, carries three rings of which two are aromatic, five rotatable bonds, and sits at cLogP 2.7. Only 8% carry defined stereochemistry, so the collection is predominantly flat, achiral scaffolding decorated with substituents.
Wide in frameworks, deep in analogs
Two measurements of the same collection pull in opposite directions, and both are true. Reduced to Bemis-Murcko scaffolds it is broad: 679 distinct frameworks per thousand molecules, 88% of them appearing exactly once, and the ten commonest together account for only 4.0%. But compared whole-molecule against whole-molecule (exhaustively, all 4,612,044 against each other with no sampling), it is dense.
The resolution is that the collection is wide in frameworks and deep in analogs around them: many scaffolds, each decorated many times over with closely related substituents. That has a direct consequence for how any model on this data must be evaluated: a random split places near-identical analogs on both sides and reports an optimistic number.
All of it is 2D structural similarity. Two compounds that look alike on paper can present different three-dimensional pharmacophores, and two that look unrelated can present similar ones. It describes the catalogue, not what the molecules do, which is the entire reason the 3D fingerprint exists.
The peptide corpus, and why it is not just more molecules
An RCSB query at 2.5 Å resolution or better with R-free at or below 0.22 returned 71,018 entries, of which 69,450 yielded usable fragments: 1,524,535 loop instances across 135,408 chains, collapsing to 117,741 distinct sequences and 133,136 distinct capped molecules. Loops are runs of two to six residues lying between annotated secondary structure, read from the helix and strand records in the mmCIF itself rather than re-derived.
Each fragment is capped with its real flanking atoms (an acetyl built from the preceding residue, an N-methylamide from the following one), so it never presents a free amine and acid the protein does not have. Backbone continuity is enforced by requiring the carbon-to-nitrogen distance between consecutive residues to fall between 1.0 and 2.0 Å rather than trusting residue numbering, which rejected 54,520 candidates that numbering alone would have accepted.
Ensemble enhancement is the novel step. Each distinct peptide gets one fingerprint that ORs together every rigid crystallographic conformation actually observed for it in the PDB with 100 computed conformers, generated exactly as for the screening collection. The conformer count therefore runs above 100, and the fingerprint carries both what nature was caught doing and what the molecule can do.
Repeats are kept rather than deduplicated, and the reason is measurable. Across 1,500 sequences seen in three or more entries, the median agreement between observations of the same sequence is 0.878, but 18% fall below 0.70, carrying genuinely different pharmacophores from one crystal structure to the next, and only 19% are near copies. Collapsing them to one representative would throw that away.
Model card
- Release
- PharmCast SP v4, trained 21 August 2026
- Architecture
- Feedforward network, 2,059 inputs to 10,560 sigmoid outputs
- Training
- 40 epochs, 232.8 minutes, 2,597,479 molecules
- Validation
- 9,975 held-out collection molecules, fixed across every release so versions stay comparable
- Median bitwise MCC
- 0.8954
- Applicability
- Catalogue chemistry to ~600 Da; peptides to ~900 Da
- Naming
- S is the screening collection, P is peptides. There is no peptide-only model.
In progress not yet released
A corpus of real, activity-backed ChEMBL molecules in the 600–1000 Da band is being fingerprinted to close exactly the gap described above; 134,765 molecules carry a measured pChEMBL value and qualify. The model that includes them will be PharmCast SCP: Screening collection, ChEMBL, Peptides. No results are claimed for it here, and none should be assumed until it is measured on the same held-out set.
Using it
The Python package is small and self-contained. words_batch is the API
that matters: the model earns its speed in a batch, and a single call is dominated by
featurisation overhead.
from pharmcast import PharmCast, pharmtan, pharmsim
pc = PharmCast.load("pharmcast_sp_v4.pt")
a, b = pc.words_batch(["CCO", "c1ccccc1O"]) # 330 uint32 each
pharmtan(a, b) # the PharmSim coefficient
pharmsim(a, b) # ...with the shared and exclusive bit sets shown
Utilities for going from 2D structure to a native fingerprint, for PharmSim / PharmTan comparison, and for expanding a fingerprint into its 0/1 bit sets will be released as an open-source package. The reference fingerprint generator is not part of that release.
Citation
A preprint describing PharmCast is in preparation. In the meantime, the method it builds on is:
McGregor, M. J. and Muskal, S. M.
Pharmacophore fingerprinting. 1.
Application to QSAR and focused library design.
J. Chem. Inf. Comput. Sci. 1999.
McGregor, M. J. and Muskal, S. M.
Pharmacophore fingerprinting. 2.
Application to primary library design.
J. Chem. Inf. Comput. Sci. 2000.