Skip to content

Bioinformatics & Multi-Omics · Data family

Data is not the bottleneck. Interpretation is.

Sequencing is cheap and analysis is commoditised. What is scarce is an interpretation you can put in front of a review committee - and that is what we build: 866,000 expression measurements, 421 span-anchored evidence records and a quoted sentence behind every claim.

866,000 expression measurementsWet-lab validation attached

Three services · analyse, discover, deploy

  1. 01Analytics
  2. 02Biomarkers
  3. 03Variants
866,000
Expression measurements, assay and unit preserved on every row
31,684
Sample size behind our primary cis-eQTL causal instruments
421
Span-anchored biomarker records across 184 targets and 86 diseases
1.00
Grounding accuracy - all 50 sampled citations resolved to a real record

The problem

A score with no sentence behind it

Institutions generate more omics data than they can analyse. Datasets sit for months, grant deliverables slip, the postdoc who knew the pipeline has left, and published biomarkers fail to reproduce because nobody validated them in the matrix that actually matters.

Most interpretation platforms answer this with a score - and a score with no sentence behind it cannot be checked, cannot be defended, and cannot be distinguished from a confident error. We built the alternative: every assertion carries the document, the character span and the quoted sentence, so a reviewer can check the claim in one click and your finding survives the committee that decides its future.

Engineering you can see

One chart we make it impossible to draw.

SourceUnitWhat it actually measures
GTExTPMBulk RNA-seq across donor tissue
Tabula SapiensCPM (pseudobulk)Single-cell RNA, summed per cell type
PRIDEPPB (iBAQ)Mass spectrometry - protein, not transcript

One shared axis · 0 – 24,000 PPB (iBAQ)

PRIDE · Protein abundance

24,000 PPB (iBAQ)

GTEx · Transcript count

35 TPM

← 0.15% of the axis. This is the whole bar.

The protein reading colours hot and the transcript reading vanishes below a pixel. Nothing here is a scaling bug - it is a category error. RNA and protein abundance correlate weakly, and "is this protein actually present in the tissue I care about?" is the exact question the panel was opened to answer.

Own scale · 0 – 24,000 PPB (iBAQ)

24,000

PPB (iBAQ)

PRIDE · Protein abundance

Own scale · 0 – 35 TPM

35

TPM

GTEx · Transcript count

Not comparable · never plotted together

Each unit is scaled within its own group and labelled with the assay that produced it. Our API groups by datasource and unit, and no chart in the platform is permitted to put two units on one scale.

The data layer

An ingestion layer we own, not a list of subscriptions

Genetics, literature, trials and a full ontology spine, normalised and linked so a question returns in seconds with every row traceable to its source. Where we chose one source over another, the reason is on the record.

SourceRole
eQTLGen (blood, N ≈ 31,684)Primary cis-eQTL instruments - the largest open cis-eQTL meta-analysis, which is what gives MR the statistical power small tissue panels lack
eQTL CatalogueTissue-specific instrument fallback
GWAS Catalog summary statisticsDisease outcome betas for two-sample MR
gnomADPopulation allele frequency, dbSNP rsID, VEP consequence, HGVSp protein change
ClinVarClinical classification, review-status star rating, associated condition, molecular consequence
Open TargetsPer-datatype evidence: genetic association, known drug, pathway, expression, literature, animal model
LayerWhat it holds
Corpus snapshotsImmutable, named states of the corpus. Every score and every generated document records the snapshot it came from, or it is not reproducible
Source documentsWritten once, content-hashed, licence-tagged
Entity mentionsWith character offsets. The atom of extraction
Evidence recordsSubject–predicate–object, with polarity, the document span that asserts it, the exact quoted sentence and a study-strength rating
Sources ingestedEurope PMC and PubMed full text · ClinicalTrials.gov protocols, results, sites, references and termination reasons kept verbatim · FDA records · patent corpus
OntologyWhat it resolves
UniProt · HGNCProtein and gene identity - the same gene means the same accession, not the same string
MONDO · EFODisease identity, so the same condition reported three ways collapses to one node
MeSHLiterature indexing vocabulary and its hierarchy
ChEBIChemical entity identity and class hierarchy

Span-level citation

Click any claim, and you land on the sentence

Span-level citation is the only citation a scientist will actually trust - and it is standard on every deliverable we produce. Each assertion carries the corpus snapshot, the source document, the character span and the exact quoted sentence.

  1. 01Corpus snapshotImmutable and named. Any past conclusion can be recomputed against the exact corpus that produced it.
  2. 02Source documentContent-hashed and licence-tagged, written once.
  3. 03Entity mentionCharacter offsets into the text. The atom of extraction.
  4. 04Evidence recordSubject–predicate–object with polarity, the asserting span and a study-strength rating.

ERBB2biomarker_forBreast carcinoma

affirmed

Asserting span · chars 1,184–1,268

Across the validation cohort, ERBB2 amplification was associated with reduced disease-free survival independently of nodal status.

Study strength
Strong · prospective cohort
Computed from
snapshot 2026-04-11

Illustrative record, in the structure the platform stores. Retracted source documents and superseded evidence rows are filtered out of live queries, not silently left in.

The three services

Analyse the data, find what survives, deploy the workflow

You can enter at any of them. Each deliverable stands on its own, and each one hands the next a starting point rather than a hand-off risk.

1Multi-omics data analyticsWhat is this data actually telling us?

The dataset you have already generated, turned into a conclusion your team can defend - and a pipeline they can re-run without us.

What you get

  • An analysis report with publication-quality figures
  • A reproducible pipeline with the full parameter record - you can run it again without us
  • The processed data package in standard formats
  • An interpretation session with your research team, because a PDF is not a conclusion

The AI capability

  • A point-in-time feature store

    One curated, versioned source of target, asset and disease features that every model reads from, instead of each analysis re-deriving the same numbers slightly differently. Reads are point-in-time: a training pull can never see a value that post-dates its label. There is an explicit leakage test in the codebase, and it runs as part of validation rather than living in a doc as an intention.

  • Knowledge-graph integration

    node2vec embeddings over the biomedical graph with a supervised edge classifier - Hadamard product of node embeddings plus common neighbours, Adamic–Adar, Jaccard, preferential attachment and random-walk-with-restart. Positives are held out before the embedding is trained, so the AUC we report is not the flattering number you get when the embedding has already walked the test edges.

  • Semantic retrieval

    Across the integrated entity corpus, so a question phrased in your words finds the record filed under someone else’s.

Measured performance

BenchmarkResult
Evidence retrieval (4 diseases, Open Targets)Macro recall@15 0.917 · provenance coverage 1.00
Semantic search (5,921-entity corpus)Retrieval accuracy@8 0.875
Target dossier generation (6 targets)Completeness 1.00 · citation coverage 1.00 across 72 cited items
Tractability profiling (812 targets)AUROC 0.98 for predicting an already-drugged target
White-space analysis (1,200 target–disease pairs)Top-decile mean evidence 0.553 vs population 0.233 · top-decile clinical precedent 0.006 vs population 0.267
2Biomarker discovery & validationWhich markers survive independent validation?

Panels with a statistical performance package behind them - and, per marker, the quoted sentence and character offset into the paper that supports it.

What you get

  • Validated biomarker panels with statistical performance packages - not candidate lists
  • Per-marker evidence with the quoted sentence and character offset into the source paper
  • An assay development route for the markers that survive
  • Independent-cohort validation design
  • Optional: full wet-lab validation and assay build through our laboratory partner

The AI capability

  • Span-anchored evidence mining

    Our biomarker layer runs on 421 span-anchored evidence records carrying the biomarker_for predicate, across 184 targets and 86 diseases - every one with a quoted sentence and a character offset into a real paper. Not a pre-computed score you have to take on faith. The evidence itself.

  • A closed 34-predicate relation schema

    Open-ended "extract any relationship" prompting produces a graph with thousands of near-synonymous edge types - inhibits, suppresses, blocks, downregulates, attenuates - that no query can traverse. Every extracted trigger maps onto one canonical predicate, and type signatures are enforced: a drug–disease pair cannot produce inhibits, because inhibits has domain asset and range target. That constraint alone removes a large class of nonsense.

  • Assertion classification

    "PRG-1042 inhibited SHP2" and "PRG-1042 did not inhibit SHP2" differ by one token and by 180 degrees. A pipeline that ignores negation builds a graph that confidently asserts the opposite of the literature, and nothing downstream can detect it because the graph looks perfectly well-formed. We run a NegEx/ConText-style scope classifier over affirmed / negated / speculated / hypothetical, deliberately rule-based so a domain expert can inspect and debug it.

  • Polarity is first-class

    Twelve papers may say X marks Y and three may say it does not. A scalar confidence value has destroyed that disagreement - and the disagreement is exactly what you need to see. We keep all fifteen records, and refuting evidence is retained as evidence.

3Clinical bioinformatics & variant interpretationHow do we report faster and more consistently?

A variant interpretation workflow deployed into your environment, with your team trained to run it after we leave.

What you get

  • A variant interpretation workflow with defined turnaround times
  • Per-variant evidence assembly: population frequency, clinical classification, predicted consequence, protein change and associated condition
  • Gene and panel-level annotation packages
  • Pipeline deployment into your environment with full parameter records
  • Training for the team who will run it after we leave

The AI capability

  • Variant evidence assembly

    ClinVar clinical classification with the review-status star rating, so a one-star single submitter is never presented as equivalent to a reviewed-by-expert-panel call - plus the associated condition, protein change and molecular consequence. gnomAD supplies dbSNP rsID, VEP consequence, HGVSp and combined exome + genome allele counts and frequency. Conditions are ontology-resolved through MONDO and EFO.

  • Formal two-sample Mendelian randomisation

    Where the question is causal rather than descriptive: cis-eQTL instruments as exposure against disease GWAS as outcome, estimated by Wald ratio, IVW, MR-Egger and weighted median, with Cochran’s Q for heterogeneity and the Egger intercept for directional pleiotropy.

  • Grounded natural-language querying

    Ask a question, get an answer with citations into the underlying records. Graph queries run from parameterised templates, never from model-written query strings - untrusted text in an ingested document must never become execution against the knowledge graph.

  • Full reproducibility

    Every generated artefact records the corpus snapshot it was computed from. Evidence is append-only in spirit: re-extraction at a new model version writes new rows and marks the old ones superseded, so any past conclusion can be reconstructed exactly.

Measured performance

BenchmarkResult
Causal-evidence classification (500 records)Accuracy 1.00 · mean 0.865 human-genetic-causal vs 0.001 limited evidence
Grounded copilot (10 questions, 50 citations checked)Grounding accuracy 1.00 · fabrication rate 0.00
Provenance coverage, all discovery benchmarks100%

Where we stop

We do not issue ACMG/AMP variant classifications as a clinical service, and we do not sign out diagnostic reports. We build, validate and deploy the workflow that assembles the evidence; the clinical call belongs to your qualified personnel under your accreditation.

Measured performance

Every gold-standard marker in breast carcinoma, recovered in the top 20

Recovery against known clinical markers across four tumour types, macro average 0.835. We publish the whole table, misses included - a vendor who only shows you the wins has not shown you a benchmark.

Breast carcinoma

1.0recall

  • BRCA1
  • BRCA2
  • ERBB2
  • ESR1
  • PIK3CA

Lung carcinoma

0.875recall

  • ALK
  • EGFR
  • ERBB2
  • KRAS
  • MET
  • RET
  • ROS1
  • BRAF

Colorectal carcinoma

0.8recall

  • BRAF
  • ERBB2
  • KRAS
  • PIK3CA
  • NRAS

Melanoma

0.667recall

  • BRAF
  • KIT
  • NRAS

Macro average

0.835

Enterprise AI for Life Sciences

Ready to Advance Your Drug Discovery Pipeline?

Partner with Prognica Labs to leverage enterprise-grade AI, computational chemistry, and molecular simulation technologies that accelerate discovery, reduce development risk, and improve R&D productivity.

From biotech startups to global pharmaceutical organizations, we help research teams make faster, evidence-driven decisions across every stage of early drug discovery.