Translational Biology · Laboratory family
The experiment that settles it
A prediction becomes knowledge at the bench and nowhere else. We design the experiment computationally, run it in the laboratory, and hand you the protocol, the data and the reasoning together.
Amplification · the moment a prediction becomes a measurement
Five services · buy one, or run them as one programme
- 866,000
- Expression measurements behind every tissue and matrix decision
- 0.98
- AUROC identifying an already-drugged target, across 812 profiles
- 31,684
- Sample size behind the cis-eQTL instruments we test causality with
- 100%
- Provenance coverage - no score without a traceable source record
The translation gap
Two vendors, one gap, and you standing in it
The vendors who generate hypotheses cannot test them. The laboratories who can test them did not generate the hypothesis and do not know what would falsify it. Between the two sits a translation gap that the client always ends up managing - usually in email, usually badly, and always at the cost of a quarter.
Somebody has to own both ends. That is what this service family is. Everything on this page is delivered with our laboratory partner: one commercial interface, one scientific plan, one timeline.
Upstream
The computational vendor
Generates the hypothesis. Cannot test it, and has no stake in whether it survives.
Downstream
The contract laboratory
Can run the experiment. Did not generate the hypothesis, and does not know what would falsify it.
You, in email, for a quarter
Re-explaining the biology, reconciling two statistical plans, and arbitrating between two parties who each believe the other one is wrong.
One plan, one team, one timeline
- Design
- Approve
- Execute
- Interpret
- Hand over
The same scientists who computed the prediction sit in the room where the experiment is designed, and the acceptance criteria are agreed in writing before anyone touches a pipette.
Before anyone touches a pipette
The bench decisions are made on data, not on habit
Which tissue. Which matrix. Which species. Which modality. Every one of those is a data question with an expensive wrong answer, and each is settled before the first reagent is ordered.
| Data | Scale / detail | Why it changes a wet-lab decision |
|---|---|---|
| Baseline expression (Open Targets) | 866,000 measurements across GTEx (TPM), Tabula Sapiens (CPM pseudobulk) and PRIDE (PPB, mass-spec protein) | Tells you which tissue and which matrix to run in - and never pools the three, because RNA and protein are different molecules with different units |
| Subcellular localisation | Per-target, with source | A target in the nucleus is not reachable by an antibody. That single fact rules out an entire modality before you spend on it |
| Ortholog identity | Species, homology type, percent identity | A mouse ortholog at 95% identity means a mouse model may transfer. At 40% it very likely does not. This is the model-organism decision, made on data |
| Gene Ontology | Function / Process / Component, kept as three separate aspects | They answer different questions and are never merged into one list |
| Tractability & safety | Small-molecule and antibody scores, best-modality tier, safety event count | Determines whether validation should be chasing a small molecule, a biologic, or neither |
| Protein identity | UniProt/SwissProt accessions, held separately from gene identity | Querying UniProt with an HGNC symbol returns nothing - conflating the two authorities is a silent data-loss bug |
| Source | What it settles |
|---|---|
| eQTLGen cis-eQTL instruments (N ≈ 31,684) | The exposure side of a formal causal test, with the statistical power small tissue panels lack |
| eQTL Catalogue | Tissue-specific instruments where blood is the wrong compartment |
| GWAS Catalog summary statistics | The disease outcome, for two-sample Mendelian randomisation |
| gnomAD | Population allele frequency and dbSNP identifiers - the variants that will break an assay in the field |
| ClinVar | Clinical classification with review-status star ratings, so a single submitter is never read as an expert panel |
| Span-cited literature | The quoted sentence and its character offset into the source paper, not a reference list |
The five services
Validate, measure, build, express - and keep the team that learned your biology
Each one is independently available and frequently bought that way. Run together, they are a single programme with one plan and one timeline behind them.
1Target validationDoes modulating this target do what we predicted?Confirm or refute
An experimental answer to the question your investment committee is actually asking - with the analysis plan agreed before the experiment runs, so the result means something whichever way it goes.
What you get
- Experimental confirmation or refutation of the target hypothesis - both are results
- Expression confirmation in the disease-relevant tissue and matrix
- Knockdown or perturbation phenotype data
- A pre-registered analysis plan agreed before the experiment runs
- Full protocols and parameter records, written for transfer to your team
What the platform contributes
Multi-evidence prioritisation with published weights
Genetic association 0.30 · known-drug precedent 0.18 · tractability 0.16 · pathway 0.12 · expression 0.10 · literature 0.08 · animal model 0.06. Every score itemises its factor contributions and traces to source records, and the weights are tunable per therapeutic strategy - a rare-disease programme should not weight genetics the way an oncology programme does. You can see what we used, and you can change it.
Formal two-sample Mendelian randomisation, before you commit bench budget
cis-eQTL instruments as exposure against disease GWAS as outcome, estimated by Wald ratio, IVW, MR-Egger and weighted median, with Cochran’s Q for heterogeneity and the Egger intercept for directional pleiotropy. A causal question gets a causal method, not a correlation dressed up as one.
Tractability tiering, read against known safety events
Clinically validated (≥0.9), structurally tractable (≥0.6), druggable family (≥0.3), challenging - so validation effort goes where a drug could actually follow it.
Knowledge-graph link prediction, honestly evaluated
node2vec embeddings with a supervised edge classifier, and test edges held out before the embedding is trained rather than after. Every predicted association returns with the evidence paths supporting it.
Ortholog-informed model selection
The species you validate in is a data decision. We make it explicitly, on percent identity and homology type, rather than defaulting to mouse and discovering the problem in month five.
Measured performance
| Benchmark | Result |
|---|---|
| Target scoring vs Open Targets (700 pairs) | Spearman ρ 0.72 · top-decile prioritisation enrichment 0.78 |
| Tractability profiles (812 targets) | AUROC 0.98 · mean 0.953 drugged vs 0.391 undrugged |
| Causal-evidence classification (500 records) | Accuracy 1.00 · 0.865 mean for human-genetic-causal vs 0.001 for limited evidence |
| Evidence retrieval (4 diseases) | Macro recall@15 0.917 · provenance coverage 1.00 |
Representative result: a five-target fibrosis shortlist reduced to two after expression confirmation and knockdown phenotyping, with one target refuted on tissue expression before any perturbation work was commissioned.
2PCR / qPCR assay developmentCan we measure this reliably, in the matrix we actually sample?Efficiency · LoD · specificity
An assay with analytical performance data behind it and a transfer package in front of it - built on a marker chosen from expression and variant data rather than from a review article.
What you get
- A designed and empirically validated PCR or qPCR assay
- Analytical performance data: efficiency, linearity, limit of detection, specificity, reproducibility
- Optimised protocol and reagent specification
- A transfer package written so your team can run it without us
- Optional progression to kit format and OEM manufacture
What the platform contributes
Transcript and target selection that keeps its assay and unit
You amplify something that is actually present in the matrix you intend to sample. Expression evidence is never pooled across RNA and protein sources, because a transcript that is measured and a protein that is present are different claims.
Variant-aware design constraints
gnomAD allele frequencies and dbSNP identifiers flag common population variants in candidate primer and probe regions. An assay sitting on a 3%-frequency SNP fails intermittently, in the sample subset that carries it - the most expensive class of assay fault there is, because it is not reproducible on demand.
ClinVar classifications with review-status star ratings
A single-submitter call is never treated as equivalent to an expert-panel one.
Ontology-resolved conditions
MONDO and EFO, so the same indication reported three ways collapses to one node instead of splitting your evidence three ways.
Span-cited literature evidence for the marker itself
The quoted sentence and its offset in the source paper. A reviewer can check the claim in one click.
Our computational contribution here is upstream of design: deciding what to measure. What follows it is empirical validation at the bench, which is the only step that ever actually settles an assay.
Representative result: a multiplex qPCR panel designed against three transcripts confirmed present by protein-level evidence, with one originally requested target dropped after expression data showed it was transcribed but not translated in the sample matrix.
3Gene synthesisWhich variants are worth synthesising at all?Sequence-verified constructs
A designed panel rather than a guessed one. The engine works on a bare sequence - no structure required - which is what makes it useful for antibodies, enzymes and de-novo designs where no structure exists.
What you get
- Synthesised, sequence-verified constructs
- Vector design and cloning to your expression system
- Variant libraries and mutant panels where the programme needs them
- Sequence-verification data and full construct records
What the platform contributes
Saturation scanning
All 19 substitutions at a chosen position, ranked rather than listed.
Stability scanning and deep mutational scanning
Ranked candidate substitutions across the sequence, or across a specified range, so the panel you order is designed against a hypothesis instead of filling a plate.
Binding-improvement suggestions at nominated positions
Where the programme has a specific interface to improve rather than a whole protein to stabilise.
ΔΔG estimation from first principles you can inspect
Change in residue volume (packing strain), hydropathy (burial mismatch) and formal charge, plus proline and glycine backbone penalties. Where no structure is supplied, per-residue burial is estimated from windowed Kyte-Doolittle hydropathy.
Codon optimisation routed to the people who own it
Synthesis strategy and codon optimisation are executed by the synthesis laboratory on their established expression-system tooling. That is their domain expertise, and we route to it rather than duplicating it badly.
Representative result: a designed 24-variant stability panel synthesised in place of a saturation library, with the ranked shortlist supplied alongside the reasoning for each position.
4Protein expressionWill this construct actually express - and can we tell before we try?Purity · yield · activity
Purified recombinant protein to an agreed specification, with the developability liabilities in the sequence found and dealt with before anyone orders a vector.
What you get
- Expressed and purified recombinant protein to agreed specification
- Expression system selection with the reasoning stated
- Purity, yield and activity data
- Constructs and protocols documented for repeat production
- Protein supplied as the reagent for your assay or structural work
What the platform contributes
Developability liability scanning, with position and severity
Asn deamidation, Asp isomerisation, acid-labile fragmentation, N-linked glycosylation sequons, free-thiol mispairing risk, Met and Trp oxidation and N-terminal pyroglutamate - each flagged at its residue, each rated, none of them buried in a paragraph.
Aggregation-propensity profiling
Windowed Kyte-Doolittle hydropathy across the sequence, with contiguous runs above threshold reported as aggregation-prone regions. Aggregation is the most common reason an expression campaign produces inclusion bodies instead of protein, and it is visible in the sequence before anyone orders a vector.
Physicochemical and structural property calculation
And, where a structure is available, secondary-structure assignment, binding-site prediction, per-residue confidence banding for predicted models, and energy minimisation.
The purpose of running this first is narrow and practical: a construct with four high-severity liabilities and an aggregation-prone region is not a construct you want to discover problems with after eight weeks of expression work.
Representative result: two of five candidate constructs deprioritised on sequence liabilities before synthesis, with the surviving construct expressing solubly at first attempt.
5Dedicated research teamsHow do we stop losing a fortnight to contracting every cycle?Reserved capacity
A named, reserved team - computational and laboratory scientists working only on your programme, accumulating the context about your biology that never appears in a statement of work but determines how good the science is by month six.
What you get
- A named, reserved team of computational and laboratory scientists
- Continuous capacity rather than repeatedly re-scoped projects
- Direct access to the scientists doing the work, not an account manager relaying questions
- Your data, constructs, protocols and models held in your tenant
- Complete methodology records, versioned, so any past result can be reconstructed
What the platform contributes
Why this model exists
Discovery is cyclical. Optimisation, validation and assay work all run in loops, and a loop broken into separately scoped projects loses a fortnight to contracting at every turn. Reserved capacity removes that overhead - and it lets a team build up context that a series of discrete projects never can.
Corpus snapshots
Every score and every generated artefact records the immutable, named corpus state it was computed from. A conclusion reached in March can be reconstructed exactly in November.
A point-in-time feature store
One versioned source of features that every model reads from, with reads guaranteed not to see values postdating their label. The leakage guard is enforced by an explicit test that runs in validation, not by an intention recorded in a document.
Append-only evidence
Re-extraction at a new model version writes new rows and marks the old ones superseded. Nothing is destroyed, so nothing has to be taken on trust.
Reserved-capacity arrangements make sense once the cadence is established. We would usually suggest starting with a single scoped project and moving to a dedicated team when the loop is obvious.
Engineering you can see
The assay that works, until it reaches the 3% of samples that break it
A primer designed on the reference sequence alone can sit directly on a common population variant. It passes validation, transfers cleanly, and then fails intermittently in the field - in exactly the samples that carry the minor allele. Switch the design mode and watch the footprint move.
Amplicon · sense strand · 60 nt shown
Intermittent failure
The 3′ end of the forward primer sits on rs11549465, minor-allele frequency 0.031. Extension is compromised on the allele that carries it, so the assay under-calls or drops out in roughly one sample in sixteen - non-reproducibly, which is what makes it so expensive to diagnose. Nothing in a standard specificity check catches this.
Clear footprint
gnomAD allele frequencies and dbSNP identifiers are constraints on the design, not a report produced after it. The footprint moves six bases upstream, clears the variant entirely, and the assay behaves the same way in every sample subset. Same primer chemistry, same validation effort, one fewer field failure.
Before you order a vector
Eighteen flagged sites and one aggregation-prone region, found in the sequence alone
This is a candidate construct scanned with the rules the engine applies: motif matching for chemical and post-translational liabilities, and a windowed Kyte-Doolittle average for aggregation propensity. Both maps below are computed from the sequence on this page, not drawn by hand.
Candidate construct · illustrative · 140 aa · 3 Cys
1
11
21
31
41
51
61
71
81
91
101
111
121
131
- High severity4Fast deamidation, isomerisation and free-thiol mispairing. Fix or justify before synthesis.
- Medium severity8Prone motifs, fragmentation and glycosylation sequons. Assess against the intended format and process.
- Low severity6Oxidation and N-terminal pyroglutamate. Usually manageable in formulation, but recorded rather than ignored.
| Position | Residues | Liability | Severity |
|---|---|---|---|
| 1 | Q | N-terminal pyroglutamate · N-terminal Q | Low |
| 28 | NG | Asn deamidation (fast) · NG | High |
| 28 | NGS | N-linked glycosylation sequon · N-X-[S/T] | Medium |
| 33 | M | Met / Trp oxidation · M · W | Low |
| 35 | W | Met / Trp oxidation · M · W | Low |
| 46 | W | Met / Trp oxidation · M · W | Low |
| 53 | DG | Asp isomerisation · DG | High |
| 60 | DS | Asp isomerisation-prone · D-[S/T/H/D] | Medium |
| 72 | NS | Asn deamidation-prone · N-[S/T/H] | Medium |
| 75 | NT | Asn deamidation-prone · N-[S/T/H] | Medium |
| 81 | M | Met / Trp oxidation · M · W | Low |
| 82 | NS | Asn deamidation-prone · N-[S/T/H] | Medium |
| 88 | DT | Asp isomerisation-prone · D-[S/T/H/D] | Medium |
| 97 | DP | Acid-labile fragmentation · DP | Medium |
| 100 | NG | Asn deamidation (fast) · NG | High |
| 100 | NGS | N-linked glycosylation sequon · N-X-[S/T] | Medium |
| 106 | W | Met / Trp oxidation · M · W | Low |
| 140 | C | Free thiol / mispairing risk - 3 cysteines · Odd cysteine count | High |
- Window
- 9 residues
- Aggregation-prone regions
- 1
- Longest run above threshold
- 18 residues
Contiguous residues above the hydropathy threshold are reported as an aggregation-prone region. Aggregation is the single most common reason an expression campaign yields inclusion bodies instead of protein - and this map costs nothing to run.
The honest boundary on the ΔΔG estimate
The stability engine behind our variant ranking is a transparent biophysical calculation - volume change, hydropathy, formal charge, proline and glycine penalties - and not a FoldX or PyRosetta free-energy computation. It is fast, inspectable and genuinely useful for ranking a synthesis panel. It is not a substitute for a rigorous ΔΔG calculation, and we do not present it as one. If your programme needs that, we will tell you.
Published weights
You can see the weights. You can change them.
Target prioritisation is a weighted sum, and a weighted sum whose weights are secret is an opinion with a decimal point on it. Ours are published, itemised per score, and tunable per therapeutic strategy.
Why we expose them
A rare-disease programme should not weight genetics the way an oncology programme does. Because the weights are exposed rather than baked in, the ranking can be argued about - which is the point of showing them.
Sums to 1.00Tunable per therapeutic strategy
- Genetic associationHuman genetics remains the strongest single predictor of clinical success.0.30
- Known-drug precedentSomebody has already shown the target can be engaged.0.18
- TractabilityWhether a drug of any modality could plausibly follow the validation.0.16
- Pathway evidenceMechanistic coherence with the biology you are treating.0.12
- ExpressionPresent in the disease-relevant tissue, in the matrix you can sample.0.10
- LiteratureSpan-cited, with the quoted sentence and its offset - never a citation count.0.08
- Animal modelWeighted lightly, and read against ortholog identity rather than assumed to transfer.0.06
Enterprise AI for Life Sciences
Ready to Advance Your Drug Discovery Pipeline?
Partner with Prognica Labs to leverage enterprise-grade AI, computational chemistry, and molecular simulation technologies that accelerate discovery, reduce development risk, and improve R&D productivity.
From biotech startups to global pharmaceutical organizations, we help research teams make faster, evidence-driven decisions across every stage of early drug discovery.
