Bioinformatics & Multi-Omics · Data family
Data is not the bottleneck. Interpretation is.
Sequencing is cheap and analysis is commoditised. What is scarce is an interpretation you can put in front of a review committee - and that is what we build: 866,000 expression measurements, 421 span-anchored evidence records and a quoted sentence behind every claim.
Three services · analyse, discover, deploy
- 866,000
- Expression measurements, assay and unit preserved on every row
- 31,684
- Sample size behind our primary cis-eQTL causal instruments
- 421
- Span-anchored biomarker records across 184 targets and 86 diseases
- 1.00
- Grounding accuracy - all 50 sampled citations resolved to a real record
The problem
A score with no sentence behind it
Institutions generate more omics data than they can analyse. Datasets sit for months, grant deliverables slip, the postdoc who knew the pipeline has left, and published biomarkers fail to reproduce because nobody validated them in the matrix that actually matters.
Most interpretation platforms answer this with a score - and a score with no sentence behind it cannot be checked, cannot be defended, and cannot be distinguished from a confident error. We built the alternative: every assertion carries the document, the character span and the quoted sentence, so a reviewer can check the claim in one click and your finding survives the committee that decides its future.
Engineering you can see
One chart we make it impossible to draw.
| Source | Unit | What it actually measures |
|---|---|---|
| GTEx | TPM | Bulk RNA-seq across donor tissue |
| Tabula Sapiens | CPM (pseudobulk) | Single-cell RNA, summed per cell type |
| PRIDE | PPB (iBAQ) | Mass spectrometry - protein, not transcript |
One shared axis · 0 – 24,000 PPB (iBAQ)
PRIDE · Protein abundance
24,000 PPB (iBAQ)
GTEx · Transcript count
35 TPM
← 0.15% of the axis. This is the whole bar.
The protein reading colours hot and the transcript reading vanishes below a pixel. Nothing here is a scaling bug - it is a category error. RNA and protein abundance correlate weakly, and "is this protein actually present in the tissue I care about?" is the exact question the panel was opened to answer.
Own scale · 0 – 24,000 PPB (iBAQ)
24,000
PPB (iBAQ)
PRIDE · Protein abundance
Own scale · 0 – 35 TPM
35
TPM
GTEx · Transcript count
Not comparable · never plotted together
Each unit is scaled within its own group and labelled with the assay that produced it. Our API groups by datasource and unit, and no chart in the platform is permitted to put two units on one scale.
The data layer
An ingestion layer we own, not a list of subscriptions
Genetics, literature, trials and a full ontology spine, normalised and linked so a question returns in seconds with every row traceable to its source. Where we chose one source over another, the reason is on the record.
| Source | Role |
|---|---|
| eQTLGen (blood, N ≈ 31,684) | Primary cis-eQTL instruments - the largest open cis-eQTL meta-analysis, which is what gives MR the statistical power small tissue panels lack |
| eQTL Catalogue | Tissue-specific instrument fallback |
| GWAS Catalog summary statistics | Disease outcome betas for two-sample MR |
| gnomAD | Population allele frequency, dbSNP rsID, VEP consequence, HGVSp protein change |
| ClinVar | Clinical classification, review-status star rating, associated condition, molecular consequence |
| Open Targets | Per-datatype evidence: genetic association, known drug, pathway, expression, literature, animal model |
| Layer | What it holds |
|---|---|
| Corpus snapshots | Immutable, named states of the corpus. Every score and every generated document records the snapshot it came from, or it is not reproducible |
| Source documents | Written once, content-hashed, licence-tagged |
| Entity mentions | With character offsets. The atom of extraction |
| Evidence records | Subject–predicate–object, with polarity, the document span that asserts it, the exact quoted sentence and a study-strength rating |
| Sources ingested | Europe PMC and PubMed full text · ClinicalTrials.gov protocols, results, sites, references and termination reasons kept verbatim · FDA records · patent corpus |
| Ontology | What it resolves |
|---|---|
| UniProt · HGNC | Protein and gene identity - the same gene means the same accession, not the same string |
| MONDO · EFO | Disease identity, so the same condition reported three ways collapses to one node |
| MeSH | Literature indexing vocabulary and its hierarchy |
| ChEBI | Chemical entity identity and class hierarchy |
Span-level citation
Click any claim, and you land on the sentence
Span-level citation is the only citation a scientist will actually trust - and it is standard on every deliverable we produce. Each assertion carries the corpus snapshot, the source document, the character span and the exact quoted sentence.
- 01Corpus snapshotImmutable and named. Any past conclusion can be recomputed against the exact corpus that produced it.
- 02Source documentContent-hashed and licence-tagged, written once.
- 03Entity mentionCharacter offsets into the text. The atom of extraction.
- 04Evidence recordSubject–predicate–object with polarity, the asserting span and a study-strength rating.
ERBB2biomarker_forBreast carcinoma
affirmedAsserting span · chars 1,184–1,268
Across the validation cohort, ERBB2 amplification was associated with reduced disease-free survival independently of nodal status.
- Study strength
- Strong · prospective cohort
- Computed from
- snapshot 2026-04-11
Illustrative record, in the structure the platform stores. Retracted source documents and superseded evidence rows are filtered out of live queries, not silently left in.
The three services
Analyse the data, find what survives, deploy the workflow
You can enter at any of them. Each deliverable stands on its own, and each one hands the next a starting point rather than a hand-off risk.
1Multi-omics data analyticsWhat is this data actually telling us?
The dataset you have already generated, turned into a conclusion your team can defend - and a pipeline they can re-run without us.
What you get
- An analysis report with publication-quality figures
- A reproducible pipeline with the full parameter record - you can run it again without us
- The processed data package in standard formats
- An interpretation session with your research team, because a PDF is not a conclusion
The AI capability
A point-in-time feature store
One curated, versioned source of target, asset and disease features that every model reads from, instead of each analysis re-deriving the same numbers slightly differently. Reads are point-in-time: a training pull can never see a value that post-dates its label. There is an explicit leakage test in the codebase, and it runs as part of validation rather than living in a doc as an intention.
Knowledge-graph integration
node2vec embeddings over the biomedical graph with a supervised edge classifier - Hadamard product of node embeddings plus common neighbours, Adamic–Adar, Jaccard, preferential attachment and random-walk-with-restart. Positives are held out before the embedding is trained, so the AUC we report is not the flattering number you get when the embedding has already walked the test edges.
Semantic retrieval
Across the integrated entity corpus, so a question phrased in your words finds the record filed under someone else’s.
Measured performance
| Benchmark | Result |
|---|---|
| Evidence retrieval (4 diseases, Open Targets) | Macro recall@15 0.917 · provenance coverage 1.00 |
| Semantic search (5,921-entity corpus) | Retrieval accuracy@8 0.875 |
| Target dossier generation (6 targets) | Completeness 1.00 · citation coverage 1.00 across 72 cited items |
| Tractability profiling (812 targets) | AUROC 0.98 for predicting an already-drugged target |
| White-space analysis (1,200 target–disease pairs) | Top-decile mean evidence 0.553 vs population 0.233 · top-decile clinical precedent 0.006 vs population 0.267 |
2Biomarker discovery & validationWhich markers survive independent validation?
Panels with a statistical performance package behind them - and, per marker, the quoted sentence and character offset into the paper that supports it.
What you get
- Validated biomarker panels with statistical performance packages - not candidate lists
- Per-marker evidence with the quoted sentence and character offset into the source paper
- An assay development route for the markers that survive
- Independent-cohort validation design
- Optional: full wet-lab validation and assay build through our laboratory partner
The AI capability
Span-anchored evidence mining
Our biomarker layer runs on 421 span-anchored evidence records carrying the biomarker_for predicate, across 184 targets and 86 diseases - every one with a quoted sentence and a character offset into a real paper. Not a pre-computed score you have to take on faith. The evidence itself.
A closed 34-predicate relation schema
Open-ended "extract any relationship" prompting produces a graph with thousands of near-synonymous edge types - inhibits, suppresses, blocks, downregulates, attenuates - that no query can traverse. Every extracted trigger maps onto one canonical predicate, and type signatures are enforced: a drug–disease pair cannot produce inhibits, because inhibits has domain asset and range target. That constraint alone removes a large class of nonsense.
Assertion classification
"PRG-1042 inhibited SHP2" and "PRG-1042 did not inhibit SHP2" differ by one token and by 180 degrees. A pipeline that ignores negation builds a graph that confidently asserts the opposite of the literature, and nothing downstream can detect it because the graph looks perfectly well-formed. We run a NegEx/ConText-style scope classifier over affirmed / negated / speculated / hypothetical, deliberately rule-based so a domain expert can inspect and debug it.
Polarity is first-class
Twelve papers may say X marks Y and three may say it does not. A scalar confidence value has destroyed that disagreement - and the disagreement is exactly what you need to see. We keep all fifteen records, and refuting evidence is retained as evidence.
3Clinical bioinformatics & variant interpretationHow do we report faster and more consistently?
A variant interpretation workflow deployed into your environment, with your team trained to run it after we leave.
What you get
- A variant interpretation workflow with defined turnaround times
- Per-variant evidence assembly: population frequency, clinical classification, predicted consequence, protein change and associated condition
- Gene and panel-level annotation packages
- Pipeline deployment into your environment with full parameter records
- Training for the team who will run it after we leave
The AI capability
Variant evidence assembly
ClinVar clinical classification with the review-status star rating, so a one-star single submitter is never presented as equivalent to a reviewed-by-expert-panel call - plus the associated condition, protein change and molecular consequence. gnomAD supplies dbSNP rsID, VEP consequence, HGVSp and combined exome + genome allele counts and frequency. Conditions are ontology-resolved through MONDO and EFO.
Formal two-sample Mendelian randomisation
Where the question is causal rather than descriptive: cis-eQTL instruments as exposure against disease GWAS as outcome, estimated by Wald ratio, IVW, MR-Egger and weighted median, with Cochran’s Q for heterogeneity and the Egger intercept for directional pleiotropy.
Grounded natural-language querying
Ask a question, get an answer with citations into the underlying records. Graph queries run from parameterised templates, never from model-written query strings - untrusted text in an ingested document must never become execution against the knowledge graph.
Full reproducibility
Every generated artefact records the corpus snapshot it was computed from. Evidence is append-only in spirit: re-extraction at a new model version writes new rows and marks the old ones superseded, so any past conclusion can be reconstructed exactly.
Measured performance
| Benchmark | Result |
|---|---|
| Causal-evidence classification (500 records) | Accuracy 1.00 · mean 0.865 human-genetic-causal vs 0.001 limited evidence |
| Grounded copilot (10 questions, 50 citations checked) | Grounding accuracy 1.00 · fabrication rate 0.00 |
| Provenance coverage, all discovery benchmarks | 100% |
Where we stop
We do not issue ACMG/AMP variant classifications as a clinical service, and we do not sign out diagnostic reports. We build, validate and deploy the workflow that assembles the evidence; the clinical call belongs to your qualified personnel under your accreditation.
Measured performance
Every gold-standard marker in breast carcinoma, recovered in the top 20
Recovery against known clinical markers across four tumour types, macro average 0.835. We publish the whole table, misses included - a vendor who only shows you the wins has not shown you a benchmark.
Breast carcinoma
1.0recall
- BRCA1
- BRCA2
- ERBB2
- ESR1
- PIK3CA
Lung carcinoma
0.875recall
- ALK
- EGFR
- ERBB2
- KRAS
- MET
- RET
- ROS1
- BRAF
Colorectal carcinoma
0.8recall
- BRAF
- ERBB2
- KRAS
- PIK3CA
- NRAS
Melanoma
0.667recall
- BRAF
- KIT
- NRAS
Macro average
0.835
Enterprise AI for Life Sciences
Ready to Advance Your Drug Discovery Pipeline?
Partner with Prognica Labs to leverage enterprise-grade AI, computational chemistry, and molecular simulation technologies that accelerate discovery, reduce development risk, and improve R&D productivity.
From biotech startups to global pharmaceutical organizations, we help research teams make faster, evidence-driven decisions across every stage of early drug discovery.
