AI Drug Discovery · Discovery family
Choose your target on evidence
A discovery engine built on 31 million compound structures, 26 validated ADMET endpoints and a benchmark for every method it runs. Every score opens up to its factors, its weights and the source record behind it.
Five stages · enter at any of them, stop at any of them
- 30.98M
- Unique compound structures we screen against, in-house
- 1.54B
- Structure-to-patent links, indexed
- 84,417
- Measured compounds training 26 validated ADMET endpoints
- 62.6×
- Mean at 1%, measured on the public DUD-E benchmark
The data layer
Most vendors query these. We ingest, normalise and link them.
This is an ingestion layer we own and re-run, not a list of subscriptions. It is why a question that would take your team a fortnight of joins returns in seconds, with every row traceable to its source.
| Source | What it gives us |
|---|---|
| Open Targets | Target–disease associations with per-datatype evidence: genetic association, known drug, affected pathway, RNA expression, literature, animal model |
| GWAS Catalog · eQTLGen · cis-eQTL panels | Instruments for formal causal testing, not just correlation |
| gnomAD · ClinVar | Population constraint and clinical variant interpretation |
| Human Protein Atlas · GTEx | Expression specificity and off-tissue liability |
| Reactome · pathway graphs | Mechanism context and network position |
| UniProt · HGNC · MONDO · EFO · MeSH · ChEBI | The ontology spine - hierarchies, synonyms and cross-references |
Six evidence classes, one ontology spine, so "the same target" always means the same accession.
| Source | Scale in our platform |
|---|---|
| Patent chemistry corpus | 30.98 million unique compound structures |
| Patent–compound occurrence map | 1.54 billion structure-to-patent-field links |
| ChEMBL | Compound registry and measured bioactivity |
| BindingDB · PubChem · Guide to Pharmacology | Affinity and pharmacology cross-reference |
| PDB (RCSB) · AlphaFold | Experimental structures and predicted models |
| PDBbind / CASF-2016 | The affinity benchmark our scoring functions are graded on |
| DUD-E | Actives and property-matched decoys for screening enrichment tests |
Two of these are benchmark sets rather than screening libraries. They are how we grade ourselves in public.
| Source | What we extract |
|---|---|
| ClinicalTrials.gov | Protocols, results, sites, references and termination reasons, captured verbatim |
| FDA records | Approval, labelling and regulatory history |
| Europe PMC · PubMed | Full-text literature with named-entity and relation extraction |
| Patent corpus | Named-entity, relation and assertion extraction over the text |
Termination reasons are kept verbatim alongside our classification of them. A compound shelved for strategy is a different opportunity from one shelved for toxicity.
ADMET capability
26 endpoints. 84,417 measured compounds. One profile per design.
Full ADMET-Tox coverage across absorption, distribution, metabolism, excretion and toxicity - grouped as the pharmacology actually is, so you see which liability you are carrying rather than a single "tox score" that hides it.
Absorption
7 endpoints
- Solubility (AqSolDB)9,980
- Lipophilicity (AZ)4,200
- PAMPA permeability2,034
- P-gp inhibition1,212
- Caco-2 permeability898
Distribution
3 endpoints
- Blood–brain barrier1,971
- Plasma protein binding1,614
- Volume of distribution1,111
Metabolism
6 endpoints
- CYP2D6 inhibition13,130
- CYP3A4 inhibition12,327
- CYP2C9 inhibition12,092
- Three CYP substrate sets
Excretion
3 endpoints
- Microsomal clearance1,102
- Hepatocyte clearance1,020
- Half-life665
Toxicity
7 endpoints
- Acute toxicity LD507,342
- Ames mutagenicity7,248
- ClinTox1,455
- hERG644
- DILI475
- Carcinogenicity278
Each model reports its training set, task type and validation metrics rather than a bare number.
The five stages
Each stage produces the input the next one needs
1Target identification & prioritisationWhich target should we commit to?
A ranked shortlist where every position on the list can be opened up and argued with - factor by factor, source record by source record.
What you get
- A ranked target list with transparent, inspectable scoring - no black boxes
- An evidence dossier per target: genetic association, expression, pathway context, perturbation data
- A druggability and tractability assessment
- A competitive and IP landscape summary
- A recommended validation plan, ready to run
The AI capability
Weighted multi-evidence scoring
Seven factors, published weights, every contribution itemised on the score. Weights are tunable per therapeutic strategy - a rare-disease programme and an oncology programme should not weight genetics identically, and you can see and change what we used.
Knowledge-graph link prediction
node2vec embeddings over the biomedical graph, then a supervised edge classifier - Hadamard product of node embeddings concatenated with common neighbours, Adamic–Adar, Jaccard, preferential attachment and random-walk-with-restart. Positives are held out before the embedding is trained, so the reported AUC is not the optimistic number you get when the embedding has already walked over the test edges.
Formal Mendelian randomisation
cis-eQTL exposure against disease GWAS outcome, with pleiotropy and heterogeneity checks - an instrument-variable causal test, reported separately from the genetics-weighted prior. When no valid instrument exists we say "insufficient" rather than inventing a verdict.
Grounded retrieval copilot
Ask in natural language, get an answer with citations to the underlying records. Graph queries run from parameterised templates, never from model-written query strings - untrusted text in an ingested patent must not become execution against the knowledge graph.
Measured performance
| Benchmark | Result |
|---|---|
| Target scoring vs Open Targets (700 pairs) | Spearman ρ 0.72 · top-decile enrichment 0.78 |
| Tractability profiles (812 targets) | AUROC 0.98 · mean 0.953 drugged vs 0.391 undrugged |
| Causal-evidence classification (500 records) | Accuracy 1.00 · 0.865 human-genetic-causal vs 0.001 limited |
| Evidence retrieval (4 diseases) | Macro recall@15 0.917 |
| Semantic search (5,921-entity corpus) | Retrieval accuracy@8 0.875 |
| Provenance coverage, all of the above | 100% |
Representative resultForty candidate targets in fibrosis narrowed to a defensible top five in seven weeks, with wet-lab confirmation of expression and knockdown phenotype for the leading two.
2Drug repurposingIs there an approved compound that already works?
Candidates that trace back through a named target bridge, filtered for safety history and commercial feasibility - so you can always answer "why this drug".
What you get
- A ranked repurposing candidate list, filtered for safety history and commercial feasibility
- A mechanistic rationale for each candidate - why it should work, not just that it scored
- A regulatory and IP pathway note per candidate
- Optional: in-vitro confirmation in a disease-relevant model
The AI capability
Target-bridged candidate generation
A drug D that hits target T, where T is strongly associated with disease X and D is not yet indicated for X, is a candidate. Score is association(T, X) × maturity(D), where maturity rewards approved and late-stage compounds because they are faster to reposition. Every candidate traces back through its target bridge.
Independent second route
Knowledge-graph drug→disease link prediction is run as a separate method, so candidates are not an artefact of one scoring approach agreeing with itself.
Shelved and discontinued asset mining
Compounds that stopped for strategic or commercial reasons are a different opportunity from compounds that stopped for toxicity. We classify termination reasons from trial registries and keep the stated reason verbatim alongside the classification.
Measured performance
| Benchmark | Result |
|---|---|
| Candidate generation (2,000 candidates) | Construct validity 1.00 · provenance 1.00 · 85.9% repurposing-ready |
| Shelved-asset classification (45-item gold set) | Accuracy 0.756 · macro F1 0.723 |
| Actionable class, same gold set | Precision 1.00 · recall 0.684 · F1 0.813 |
The shelved-asset classifier is precision-heavy by design: when it says an asset is actionable it has not yet been wrong on our gold set, and it accepts missing about a third of them to keep that true. For a repurposing programme, a false positive burns a diligence cycle; a false negative adds one line to a longer list.
Representative resultThree approved compounds identified with mechanistic rationale in a rare neuromuscular disorder; two confirmed active in patient-derived cells.
3Virtual screening & hit discoveryWhat should we test first?
A staged cascade over libraries up to the full 31-million-compound corpus, ending in a diversity-selected list your chemists can actually buy.
What you get
- A prioritised, diversity-selected hit list with predicted binding modes
- The full screening methodology and parameter record
- A compound sourcing plan with availability and lead times
- Optional: confirmatory biochemical or cell-based assay results
The AI capability
Ligand-based screening
Morgan-fingerprint similarity against reference actives, with 3-D shape similarity added on the survivors. Lipinski and property filters, PAINS and Brenk structural alerts, then diversity selection so you do not buy twenty analogues of the same scaffold.
Structure-based docking
Pocket detection, pharmacophore definition, a Vina-like scoring function and physics-based rescoring of surviving poses.
QSAR models
Trained on ChEMBL target activity or Tox21 endpoints - classification with calibrated probabilities and regression, both on scaffold-aware splits, because a random split on congeneric series flatters every model ever built.
Measured performance
| Benchmark | Result |
|---|---|
| Descriptor fidelity (247 molecules, vs ChEMBL) | MW MAE 0.0 · logP Spearman 0.927 · Rule-of-5 exact match 89.1% |
| Structural alerts, known-liability panel | 100% recall |
| Structural alerts, approved drugs | 0% false-positive rate |
Representative resultSix million commercially available compounds screened against a kinase target, delivering 120 diversity-selected hits, of which 14 confirmed activity in a biochemical assay.
4Molecular dynamics & free-energy simulationWhich series deserves the synthesis budget?
Atomistic simulation on OpenMM with a full analysis stack - because a pose that scores well and dissociates in two nanoseconds should not outrank one that holds.
What you get
- A simulation report with trajectory and interaction-stability analysis
- Relative binding free-energy ranking across your series
- Interaction and stability maps
- Structure-based design recommendations for the next chemistry cycle
- The trajectory data package, for your own reanalysis
The AI capability
Stability
Backbone RMSD, radius of gyration and per-residue fluctuation across the trajectory.
Collective motion
Principal component analysis of the Cα trajectory and dynamic cross-correlation maps - how the pocket moves as a unit, and what moves with it.
Interactions
Residue contact maps, hydrogen-bond occupancy over the trajectory and salt-bridge persistence.
Burial
Per-frame solvent-accessible surface area.
Free-energy landscape
Projected onto the leading principal components with RMSD and radius of gyration - the basins the system actually occupies.
Binding energy
Endpoint scoring with an SASA-derived nonpolar term, combined with MD stability signals into a composite that penalises a good score achieved in a pose that does not hold.
Quantum descriptors
Extended-Hückel HOMO/LUMO, dipole moment and derived reactivity descriptors where electronic structure is the question.
Reporting binding energy without stability is how programmes lose a year to a series that was never going to hold.
Representative resultFour chemical series ranked by free-energy calculation, redirecting a medicinal chemistry programme away from a series that later proved unstable in the pocket.
5Lead optimisation & ADMETHow do we make this a drug, not just a binder?
A Pareto front across potency, ADMET and synthesisability - not a single compound from a scalarised score that hid the trade-off you cared about.
What you get
- Optimised compound designs with predicted property profiles
- An ADMET and liability risk report, including the liabilities you did not ask about
- Structure–activity relationship analysis across your series
- A synthesis prioritisation list your chemists can act on
- Optional: in-vitro ADME confirmation of the selected compounds
The AI capability
Generative design under multiple objectives
An evolutionary search over chemical space - crossover and mutation on validated structures, with an RDKit validity gate on every generated molecule. Multi-objective selection uses non-dominated sorting with crowding distance, so you receive a Pareto front rather than one compound from a hidden trade-off.
Matched molecular pairs & activity cliffs
Which single transformation moved potency, and where the series is fragile - the SAR your chemists can reason with.
Multi-parameter optimisation scoring
Desirability functions per property, so "too lipophilic" and "too polar" both count against a design, which a linear score cannot express.
Synthetic accessibility & retrosynthesis
Every design carries an accessibility band and a proposed route to building blocks. A design your chemists cannot make is not a design.
ADMET prediction across 26 endpoints
Trained on 84,417 measured compounds covering the full ADMET-Tox span - absorption, distribution, metabolism, excretion and toxicity - each model reporting its training set, task type and validation metrics.
Liability screening
PAINS and Brenk structural-alert catalogues: 100% recall on known problem chemotypes, zero false positives on approved drugs.
Representative resultFour optimisation cycles improving predicted oral bioavailability and removing an hERG liability while holding sub-100 nM potency.
Inside the target score
Seven factors. Published weights. Tuned to your therapeutic strategy.
This is the whole scoring model, not a summary of it - and it is yours to configure. A rare-disease programme and an oncology programme should not weight genetics identically, so the weights move, and the set we used is written into your report.
Reported separately
Mendelian randomisation is reported separately from this prior, not folded into it. A weighted score and a causal test answer different questions and should not be averaged together.
Genetic association
GWAS, rare-variant burden and constraint
0.30
Known-drug precedent
Existing chemical matter against the target
0.18
Target tractability
Pocket, modality and structural feasibility
0.16
Pathway evidence
Mechanism context and network position
0.12
Expression evidence
Tissue specificity and off-tissue liability
0.10
Literature evidence
Extracted assertions, weighted by study type
0.08
Animal-model evidence
Perturbation phenotypes and translatability
0.06
Screening enrichment
62.6× mean enrichment at 1%, on the open DUD-E benchmark
An enrichment factor of 62× means the top percentile of a screened library is enriched sixty-two-fold for real actives against random selection. Millions of compounds triaged to a bench-ready shortlist - and the decoy counts are published, so you can check the claim.
EGFR
- Actives
- 542
- Decoys
- 35,050
- AUROC
- 0.975
- EF 5%
- 18.4×
68.8× EF 1%
Factor Xa
- Actives
- 537
- Decoys
- 28,325
- AUROC
- 0.973
- EF 5%
- 18×
56.3× EF 1%
Mean
- Actives
- —
- Decoys
- —
- AUROC
- 0.974
- EF 5%
- 18.2×
62.6× EF 1%
Enterprise AI for Life Sciences
Ready to Advance Your Drug Discovery Pipeline?
Partner with Prognica Labs to leverage enterprise-grade AI, computational chemistry, and molecular simulation technologies that accelerate discovery, reduce development risk, and improve R&D productivity.
From biotech startups to global pharmaceutical organizations, we help research teams make faster, evidence-driven decisions across every stage of early drug discovery.
