Skip to content

Translational Biology · Laboratory family

The experiment that settles it

A prediction becomes knowledge at the bench and nowhere else. We design the experiment computationally, run it in the laboratory, and hand you the protocol, the data and the reasoning together.

Computational design + laboratory execution on one contract
866,000
Expression measurements behind every tissue and matrix decision
0.98
AUROC identifying an already-drugged target, across 812 profiles
31,684
Sample size behind the cis-eQTL instruments we test causality with
100%
Provenance coverage - no score without a traceable source record

The translation gap

Two vendors, one gap, and you standing in it

The vendors who generate hypotheses cannot test them. The laboratories who can test them did not generate the hypothesis and do not know what would falsify it. Between the two sits a translation gap that the client always ends up managing - usually in email, usually badly, and always at the cost of a quarter.

Somebody has to own both ends. That is what this service family is. Everything on this page is delivered with our laboratory partner: one commercial interface, one scientific plan, one timeline.

Upstream

The computational vendor

Generates the hypothesis. Cannot test it, and has no stake in whether it survives.

Downstream

The contract laboratory

Can run the experiment. Did not generate the hypothesis, and does not know what would falsify it.

You, in email, for a quarter

Re-explaining the biology, reconciling two statistical plans, and arbitrating between two parties who each believe the other one is wrong.

One plan, one team, one timeline

  1. Design
  2. Approve
  3. Execute
  4. Interpret
  5. Hand over

The same scientists who computed the prediction sit in the room where the experiment is designed, and the acceptance criteria are agreed in writing before anyone touches a pipette.

Before anyone touches a pipette

The bench decisions are made on data, not on habit

Which tissue. Which matrix. Which species. Which modality. Every one of those is a data question with an expensive wrong answer, and each is settled before the first reagent is ordered.

DataScale / detailWhy it changes a wet-lab decision
Baseline expression (Open Targets)866,000 measurements across GTEx (TPM), Tabula Sapiens (CPM pseudobulk) and PRIDE (PPB, mass-spec protein)Tells you which tissue and which matrix to run in - and never pools the three, because RNA and protein are different molecules with different units
Subcellular localisationPer-target, with sourceA target in the nucleus is not reachable by an antibody. That single fact rules out an entire modality before you spend on it
Ortholog identitySpecies, homology type, percent identityA mouse ortholog at 95% identity means a mouse model may transfer. At 40% it very likely does not. This is the model-organism decision, made on data
Gene OntologyFunction / Process / Component, kept as three separate aspectsThey answer different questions and are never merged into one list
Tractability & safetySmall-molecule and antibody scores, best-modality tier, safety event countDetermines whether validation should be chasing a small molecule, a biologic, or neither
Protein identityUniProt/SwissProt accessions, held separately from gene identityQuerying UniProt with an HGNC symbol returns nothing - conflating the two authorities is a silent data-loss bug
SourceWhat it settles
eQTLGen cis-eQTL instruments (N ≈ 31,684)The exposure side of a formal causal test, with the statistical power small tissue panels lack
eQTL CatalogueTissue-specific instruments where blood is the wrong compartment
GWAS Catalog summary statisticsThe disease outcome, for two-sample Mendelian randomisation
gnomADPopulation allele frequency and dbSNP identifiers - the variants that will break an assay in the field
ClinVarClinical classification with review-status star ratings, so a single submitter is never read as an expert panel
Span-cited literatureThe quoted sentence and its character offset into the source paper, not a reference list

The five services

Validate, measure, build, express - and keep the team that learned your biology

Each one is independently available and frequently bought that way. Run together, they are a single programme with one plan and one timeline behind them.

1Target validationDoes modulating this target do what we predicted?Confirm or refute

An experimental answer to the question your investment committee is actually asking - with the analysis plan agreed before the experiment runs, so the result means something whichever way it goes.

What you get

  • Experimental confirmation or refutation of the target hypothesis - both are results
  • Expression confirmation in the disease-relevant tissue and matrix
  • Knockdown or perturbation phenotype data
  • A pre-registered analysis plan agreed before the experiment runs
  • Full protocols and parameter records, written for transfer to your team

What the platform contributes

  • Multi-evidence prioritisation with published weights

    Genetic association 0.30 · known-drug precedent 0.18 · tractability 0.16 · pathway 0.12 · expression 0.10 · literature 0.08 · animal model 0.06. Every score itemises its factor contributions and traces to source records, and the weights are tunable per therapeutic strategy - a rare-disease programme should not weight genetics the way an oncology programme does. You can see what we used, and you can change it.

  • Formal two-sample Mendelian randomisation, before you commit bench budget

    cis-eQTL instruments as exposure against disease GWAS as outcome, estimated by Wald ratio, IVW, MR-Egger and weighted median, with Cochran’s Q for heterogeneity and the Egger intercept for directional pleiotropy. A causal question gets a causal method, not a correlation dressed up as one.

  • Tractability tiering, read against known safety events

    Clinically validated (≥0.9), structurally tractable (≥0.6), druggable family (≥0.3), challenging - so validation effort goes where a drug could actually follow it.

  • Knowledge-graph link prediction, honestly evaluated

    node2vec embeddings with a supervised edge classifier, and test edges held out before the embedding is trained rather than after. Every predicted association returns with the evidence paths supporting it.

  • Ortholog-informed model selection

    The species you validate in is a data decision. We make it explicitly, on percent identity and homology type, rather than defaulting to mouse and discovering the problem in month five.

Measured performance

BenchmarkResult
Target scoring vs Open Targets (700 pairs)Spearman ρ 0.72 · top-decile prioritisation enrichment 0.78
Tractability profiles (812 targets)AUROC 0.98 · mean 0.953 drugged vs 0.391 undrugged
Causal-evidence classification (500 records)Accuracy 1.00 · 0.865 mean for human-genetic-causal vs 0.001 for limited evidence
Evidence retrieval (4 diseases)Macro recall@15 0.917 · provenance coverage 1.00

Representative result: a five-target fibrosis shortlist reduced to two after expression confirmation and knockdown phenotyping, with one target refuted on tissue expression before any perturbation work was commissioned.

2PCR / qPCR assay developmentCan we measure this reliably, in the matrix we actually sample?Efficiency · LoD · specificity

An assay with analytical performance data behind it and a transfer package in front of it - built on a marker chosen from expression and variant data rather than from a review article.

What you get

  • A designed and empirically validated PCR or qPCR assay
  • Analytical performance data: efficiency, linearity, limit of detection, specificity, reproducibility
  • Optimised protocol and reagent specification
  • A transfer package written so your team can run it without us
  • Optional progression to kit format and OEM manufacture

What the platform contributes

  • Transcript and target selection that keeps its assay and unit

    You amplify something that is actually present in the matrix you intend to sample. Expression evidence is never pooled across RNA and protein sources, because a transcript that is measured and a protein that is present are different claims.

  • Variant-aware design constraints

    gnomAD allele frequencies and dbSNP identifiers flag common population variants in candidate primer and probe regions. An assay sitting on a 3%-frequency SNP fails intermittently, in the sample subset that carries it - the most expensive class of assay fault there is, because it is not reproducible on demand.

  • ClinVar classifications with review-status star ratings

    A single-submitter call is never treated as equivalent to an expert-panel one.

  • Ontology-resolved conditions

    MONDO and EFO, so the same indication reported three ways collapses to one node instead of splitting your evidence three ways.

  • Span-cited literature evidence for the marker itself

    The quoted sentence and its offset in the source paper. A reviewer can check the claim in one click.

Our computational contribution here is upstream of design: deciding what to measure. What follows it is empirical validation at the bench, which is the only step that ever actually settles an assay.

Representative result: a multiplex qPCR panel designed against three transcripts confirmed present by protein-level evidence, with one originally requested target dropped after expression data showed it was transcribed but not translated in the sample matrix.

3Gene synthesisWhich variants are worth synthesising at all?Sequence-verified constructs

A designed panel rather than a guessed one. The engine works on a bare sequence - no structure required - which is what makes it useful for antibodies, enzymes and de-novo designs where no structure exists.

What you get

  • Synthesised, sequence-verified constructs
  • Vector design and cloning to your expression system
  • Variant libraries and mutant panels where the programme needs them
  • Sequence-verification data and full construct records

What the platform contributes

  • Saturation scanning

    All 19 substitutions at a chosen position, ranked rather than listed.

  • Stability scanning and deep mutational scanning

    Ranked candidate substitutions across the sequence, or across a specified range, so the panel you order is designed against a hypothesis instead of filling a plate.

  • Binding-improvement suggestions at nominated positions

    Where the programme has a specific interface to improve rather than a whole protein to stabilise.

  • ΔΔG estimation from first principles you can inspect

    Change in residue volume (packing strain), hydropathy (burial mismatch) and formal charge, plus proline and glycine backbone penalties. Where no structure is supplied, per-residue burial is estimated from windowed Kyte-Doolittle hydropathy.

  • Codon optimisation routed to the people who own it

    Synthesis strategy and codon optimisation are executed by the synthesis laboratory on their established expression-system tooling. That is their domain expertise, and we route to it rather than duplicating it badly.

Representative result: a designed 24-variant stability panel synthesised in place of a saturation library, with the ranked shortlist supplied alongside the reasoning for each position.

4Protein expressionWill this construct actually express - and can we tell before we try?Purity · yield · activity

Purified recombinant protein to an agreed specification, with the developability liabilities in the sequence found and dealt with before anyone orders a vector.

What you get

  • Expressed and purified recombinant protein to agreed specification
  • Expression system selection with the reasoning stated
  • Purity, yield and activity data
  • Constructs and protocols documented for repeat production
  • Protein supplied as the reagent for your assay or structural work

What the platform contributes

  • Developability liability scanning, with position and severity

    Asn deamidation, Asp isomerisation, acid-labile fragmentation, N-linked glycosylation sequons, free-thiol mispairing risk, Met and Trp oxidation and N-terminal pyroglutamate - each flagged at its residue, each rated, none of them buried in a paragraph.

  • Aggregation-propensity profiling

    Windowed Kyte-Doolittle hydropathy across the sequence, with contiguous runs above threshold reported as aggregation-prone regions. Aggregation is the most common reason an expression campaign produces inclusion bodies instead of protein, and it is visible in the sequence before anyone orders a vector.

  • Physicochemical and structural property calculation

    And, where a structure is available, secondary-structure assignment, binding-site prediction, per-residue confidence banding for predicted models, and energy minimisation.

The purpose of running this first is narrow and practical: a construct with four high-severity liabilities and an aggregation-prone region is not a construct you want to discover problems with after eight weeks of expression work.

Representative result: two of five candidate constructs deprioritised on sequence liabilities before synthesis, with the surviving construct expressing solubly at first attempt.

5Dedicated research teamsHow do we stop losing a fortnight to contracting every cycle?Reserved capacity

A named, reserved team - computational and laboratory scientists working only on your programme, accumulating the context about your biology that never appears in a statement of work but determines how good the science is by month six.

What you get

  • A named, reserved team of computational and laboratory scientists
  • Continuous capacity rather than repeatedly re-scoped projects
  • Direct access to the scientists doing the work, not an account manager relaying questions
  • Your data, constructs, protocols and models held in your tenant
  • Complete methodology records, versioned, so any past result can be reconstructed

What the platform contributes

  • Why this model exists

    Discovery is cyclical. Optimisation, validation and assay work all run in loops, and a loop broken into separately scoped projects loses a fortnight to contracting at every turn. Reserved capacity removes that overhead - and it lets a team build up context that a series of discrete projects never can.

  • Corpus snapshots

    Every score and every generated artefact records the immutable, named corpus state it was computed from. A conclusion reached in March can be reconstructed exactly in November.

  • A point-in-time feature store

    One versioned source of features that every model reads from, with reads guaranteed not to see values postdating their label. The leakage guard is enforced by an explicit test that runs in validation, not by an intention recorded in a document.

  • Append-only evidence

    Re-extraction at a new model version writes new rows and marks the old ones superseded. Nothing is destroyed, so nothing has to be taken on trust.

Reserved-capacity arrangements make sense once the cadence is established. We would usually suggest starting with a single scoped project and moving to a dedicated team when the loop is obvious.

Engineering you can see

The assay that works, until it reaches the 3% of samples that break it

A primer designed on the reference sequence alone can sit directly on a common population variant. It passes validation, transfers cleanly, and then fails intermittently in the field - in exactly the samples that carry the minor allele. Switch the design mode and watch the footprint move.

Amplicon · sense strand · 60 nt shown

Forward primer 5′→3′
GCTAAGCCTGACTTCAGGTCAACGTGGCATCAGCTGAAGCCTAGGTCATGCAACGTTGAC
rs11549465 · C>T · gnomAD AF 0.031

Intermittent failure

The 3′ end of the forward primer sits on rs11549465, minor-allele frequency 0.031. Extension is compromised on the allele that carries it, so the assay under-calls or drops out in roughly one sample in sixteen - non-reproducibly, which is what makes it so expensive to diagnose. Nothing in a standard specificity check catches this.

Clear footprint

gnomAD allele frequencies and dbSNP identifiers are constraints on the design, not a report produced after it. The footprint moves six bases upstream, clears the variant entirely, and the assay behaves the same way in every sample subset. Same primer chemistry, same validation effort, one fewer field failure.

Before you order a vector

Eighteen flagged sites and one aggregation-prone region, found in the sequence alone

This is a candidate construct scanned with the rules the engine applies: motif matching for chemical and post-translational liabilities, and a windowed Kyte-Doolittle average for aggregation propensity. Both maps below are computed from the sequence on this page, not drawn by hand.

Candidate construct · illustrative · 140 aa · 3 Cys

1

QQ1 · lowN-terminal pyroglutamateVQLVESGGG

11

LVQPGGSLRL

21

SCAASGFNN28 · highAsn deamidation (fast) · N-linked glycosylation sequonGG29 · highAsn deamidation (fast) · N-linked glycosylation sequonSS30 · mediumN-linked glycosylation sequon

31

YAMM33 · lowMet / Trp oxidationSWW35 · lowMet / Trp oxidationVRQAP

41

GKGLEWW46 · lowMet / Trp oxidationVSAI

51

SGDD53 · highAsp isomerisationGG54 · highAsp isomerisationSTYYADD60 · mediumAsp isomerisation-prone

61

SS61 · mediumAsp isomerisation-proneVKGRFTISR

71

DNN72 · mediumAsn deamidation-proneSS73 · mediumAsn deamidation-proneKNN75 · mediumAsn deamidation-proneTT76 · mediumAsn deamidation-proneLYLQ

81

MM81 · lowMet / Trp oxidationNN82 · mediumAsn deamidation-proneSS83 · mediumAsn deamidation-proneLRAEDD88 · mediumAsp isomerisation-proneTT89 · mediumAsp isomerisation-proneA

91

VYYCARDD97 · mediumAcid-labile fragmentationPP98 · mediumAcid-labile fragmentationGNN100 · highAsn deamidation (fast) · N-linked glycosylation sequon

101

GG101 · highAsn deamidation (fast) · N-linked glycosylation sequonSS102 · mediumN-linked glycosylation sequonLDYWW106 · lowMet / Trp oxidationGQGT

111

LVTVSSAVLA

121

VLLIVFAGVA

131

LTKGPSVFPCC140 · highFree thiol / mispairing risk - 3 cysteines
  • High severity4Fast deamidation, isomerisation and free-thiol mispairing. Fix or justify before synthesis.
  • Medium severity8Prone motifs, fragmentation and glycosylation sequons. Assess against the intended format and process.
  • Low severity6Oxidation and N-terminal pyroglutamate. Usually manageable in formulation, but recorded rather than ignored.
PositionResiduesLiabilitySeverity
1QN-terminal pyroglutamate · N-terminal QLow
28NGAsn deamidation (fast) · NGHigh
28NGSN-linked glycosylation sequon · N-X-[S/T]Medium
33MMet / Trp oxidation · M · WLow
35WMet / Trp oxidation · M · WLow
46WMet / Trp oxidation · M · WLow
53DGAsp isomerisation · DGHigh
60DSAsp isomerisation-prone · D-[S/T/H/D]Medium
72NSAsn deamidation-prone · N-[S/T/H]Medium
75NTAsn deamidation-prone · N-[S/T/H]Medium
81MMet / Trp oxidation · M · WLow
82NSAsn deamidation-prone · N-[S/T/H]Medium
88DTAsp isomerisation-prone · D-[S/T/H/D]Medium
97DPAcid-labile fragmentation · DPMedium
100NGAsn deamidation (fast) · NGHigh
100NGSN-linked glycosylation sequon · N-X-[S/T]Medium
106WMet / Trp oxidation · M · WLow
140CFree thiol / mispairing risk - 3 cysteines · Odd cysteine countHigh
APR 113–130 · peak 3.68threshold 1.04.1-2.4residue 1140
Window
9 residues
Aggregation-prone regions
1
Longest run above threshold
18 residues

Contiguous residues above the hydropathy threshold are reported as an aggregation-prone region. Aggregation is the single most common reason an expression campaign yields inclusion bodies instead of protein - and this map costs nothing to run.

The honest boundary on the ΔΔG estimate

The stability engine behind our variant ranking is a transparent biophysical calculation - volume change, hydropathy, formal charge, proline and glycine penalties - and not a FoldX or PyRosetta free-energy computation. It is fast, inspectable and genuinely useful for ranking a synthesis panel. It is not a substitute for a rigorous ΔΔG calculation, and we do not present it as one. If your programme needs that, we will tell you.

Published weights

You can see the weights. You can change them.

Target prioritisation is a weighted sum, and a weighted sum whose weights are secret is an opinion with a decimal point on it. Ours are published, itemised per score, and tunable per therapeutic strategy.

Why we expose them

A rare-disease programme should not weight genetics the way an oncology programme does. Because the weights are exposed rather than baked in, the ranking can be argued about - which is the point of showing them.

Sums to 1.00Tunable per therapeutic strategy

  • Genetic associationHuman genetics remains the strongest single predictor of clinical success.0.30
  • Known-drug precedentSomebody has already shown the target can be engaged.0.18
  • TractabilityWhether a drug of any modality could plausibly follow the validation.0.16
  • Pathway evidenceMechanistic coherence with the biology you are treating.0.12
  • ExpressionPresent in the disease-relevant tissue, in the matrix you can sample.0.10
  • LiteratureSpan-cited, with the quoted sentence and its offset - never a citation count.0.08
  • Animal modelWeighted lightly, and read against ortholog identity rather than assumed to transfer.0.06

Enterprise AI for Life Sciences

Ready to Advance Your Drug Discovery Pipeline?

Partner with Prognica Labs to leverage enterprise-grade AI, computational chemistry, and molecular simulation technologies that accelerate discovery, reduce development risk, and improve R&D productivity.

From biotech startups to global pharmaceutical organizations, we help research teams make faster, evidence-driven decisions across every stage of early drug discovery.