Glossary · Reproducibility and Workflows
Multi-omics integration
Merging RNA-seq, proteomics, and methylation into one model only works if you pick the method for the question you're actually asking.
By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed September 2026 · 3 min read
Also: MOFA, multi-modal integration
Definition
Multi-omics integration is the joint statistical analysis of two or more molecular datasets collected on the same or overlapping samples, such as genomics, transcriptomics, proteomics, methylation, and metabolomics, done to reveal patterns that no single omics layer shows alone. It is not concatenation: each modality has its own scale, sparsity, and noise profile, so integration methods normalize or model those differences before extracting shared signal. Methods fall into three integration strategies (early, intermediate, late) and three method families (statistical, such as WGCNA and SNF; multivariate/latent-variable, such as MOFA2 and DIABLO; and machine learning). The two most common tools split by goal: MOFA2 is unsupervised factor discovery, DIABLO is supervised biomarker selection.
You reach for multi-omics integration when a single assay leaves a gap you can see but can't explain: RNA-seq shows a pathway is dysregulated, but you don't know whether that's driven by copy number, methylation, or protein turnover. The term covers a family of methods, not one tool, and the method you pick has to match the question, not the data type you happen to have most of.
This decision sits at the start of the analysis, not the end. Before you touch a package, you need to know whether you're hunting for shared structure across omics layers with no label (unsupervised, MOFA2's job) or trying to find features that predict a known outcome like disease status or drug response (supervised, DIABLO's job). Get that wrong and every downstream plot looks plausible and means nothing.
Why it matters
Pick the wrong integration strategy and you get a result that looks clean and is wrong. A chronic kidney disease study (n=37) integrated tissue transcriptomics, urine proteomics, plasma proteomics, and targeted metabolomics using both MOFA2 and DIABLO because the two methods answer different questions: MOFA2 surfaced which molecular layers co-varied, DIABLO narrowed that down to a sparse, predictive signature of 8 urinary proteins and 3 enriched pathways (complement/coagulation, cytokine signaling, JAK/STAT). That signature only became usable because the researchers validated it in an independent cohort of 94 additional patients, where it held up even after adjusting for clinical variables.
Skip validation and you're publishing correlation as if it were mechanism. Skip the strategy choice and you risk the more common failure: naive concatenation of unbalanced matrices (200 RNA-seq samples, 150 proteomics, 180 methylation) that either drowns real biological signal in the noisiest modality or manufactures clusters that track missing-data patterns and batch effects instead of biology.
Where people get it wrong
The most common mistake is treating "integration" as a data-wrangling step, merge the matrices, run PCA, done, rather than a modeling choice. That skips the normalization each modality needs (log2 transform for count data, z-scoring within each omics layer, filtering low-abundance features) and skips the strategy decision (early integration inflates redundancy, late integration misses inter-omics correlations, intermediate sits between the two).
A second, subtler mistake is picking MOFA2 when you actually want DIABLO, or the reverse. MOFA2 has no concept of "disease" or "outcome"; it finds latent factors of variation and you interpret them after the fact. DIABLO is supervised from the start and will happily find a discriminative signature even on noise if you don't hold out a validation set. Practitioners who don't separate the two end up either over-interpreting an unsupervised factor as a biomarker, or under-trusting a supervised signature because they benchmarked it against the wrong question. Both MOFA2 and DIABLO find correlated structure, not causal structure; a factor loading heavily on a pathway is a hypothesis, not a mechanism, until it's checked against known biology or an independent cohort.
A concrete example
Before reaching for MOFA2 or DIABLO, get comfortable with the matrix factorization these methods build on. Non-negative Matrix Factorization (NMF) decomposes a single omics matrix X (genes x samples) into a basis matrix W (genes x latent factors) and a coefficient matrix H (latent factors x samples), which is the same logic MOFA2 extends across multiple omics matrices at once. Run it on a 20,000-gene RNA-seq matrix asking for 5 factors, and each factor in W is a candidate hidden pathway; each column of H tells you how much that pathway is active in each sample. NMF is the right starting point because gene expression and protein abundance are non-negative by nature, so the factorization stays biologically interpretable, unlike PCA components which can go negative.
library(NMF)
res <- nmf(expr_matrix, 5)
basis_matrix <- basis(res) # genes x 5 latent factors
coef_matrix <- coef(res) # 5 latent factors x samplesRelated terms
- Single-cell integration
- Matrix factorization
- Batch effect correction
- Dimensionality reduction
- Non-negative matrix factorization
Questions people ask
- What is multi-omics integration?
It's the joint analysis of two or more molecular datasets, such as genomics, transcriptomics, proteomics, and metabolomics, collected on the same or overlapping samples, done to find patterns across layers that no single dataset shows on its own. It requires normalizing each modality to its own scale before combining them, not just pasting matrices together.
- What is the difference between MOFA2 and DIABLO?
MOFA2 is unsupervised: it finds latent factors of variation shared across omics layers without using any outcome label, similar to PCA generalized to multiple data types. DIABLO is supervised: it selects a sparse set of features across omics layers that best discriminate a known outcome, such as disease versus control. Pick MOFA2 to discover structure, DIABLO to build a predictive signature.
- Is a MOFA2 tutorial enough to run multi-omics integration correctly?
A tutorial teaches the syntax, not the judgment call that precedes it: whether your question is unsupervised discovery or supervised prediction, and whether your omics matrices are complete enough to integrate without creating phantom clusters from missing samples. Run the tutorial, but decide the strategy (early, intermediate, or late integration) and validate any resulting signature in an independent cohort before trusting it.
- Why does multi-omics integration find false clusters?
When sample coverage differs across omics layers (RNA-seq on 200 samples, proteomics on 150, methylation on 180) and you concatenate everything naively, missing-data patterns and batch effects can dominate the signal and produce clusters that track technical artifacts instead of biology. Filtering, normalizing each layer separately, and checking whether a cluster survives when you drop one omics layer catches this.
- Can multi-omics integration prove causation?
No. These methods surface correlated or co-varying signal across molecular layers, not causal relationships. Treat any factor or biomarker signature as a hypothesis that needs validation against known pathways, an orthogonal experiment, or an independent patient cohort before you act on it.
Related pages
Related reading on the blog
Sources
- Leveraging complementary multi-omics data integration methods for mechanistic insights in kidney diseases — Kidney disease case study using both MOFA2 and DIABLO, with independent-cohort validation of the resulting biomarker signature
- Algorithms and tools for data-driven omics integration to achieve multilayer biological insights: a narrative review — Categorization of integration methods into statistical, multivariate/latent-variable, and machine learning approaches
- Protocol for integrating and interpreting multi-omics data combining unsupervised and supervised data integrating approaches — Early, intermediate, and late integration strategies and standard preprocessing steps
- MOFA2: Multi-Omics Factor Analysis v2 (Bioconductor) — MOFA2 package description, version, dependencies, and the MEFISTO temporal/spatial extension