The omics revolution produces data faster than it produces understanding. Turning a mass spectrometer's output into a biological claim takes a chain of decisions β how to process, filter, normalise, test, annotate and integrate β and each link in that chain can change the answer.
This three-day intensive course offers practical, up-to-date training in data science applied to omics data, with a focus on mass-spectrometry proteomics and metabolomics and on multi-omics integration.
The aim is to train researchers, bioinformaticians and health-science professionals to manage the full path from raw instrument files to interpretable biology: processing techniques, statistical analysis, visualisation, functional enrichment, integration of several omics layers, and the application and visualisation of biological networks derived from them.
What makes the course concrete is that it follows one real cohort from start to finish. Every notebook, every plot and every exercise uses serum proteomics and metabolomics from the same 45 septic patients, so by the end of Day 3 the class has built a complete multi-omics story β including its uncertainties. All practical sessions are in Python, in Jupyter notebooks that run on Google Colab; no local installation is required.
Proteomics, metabolomics, multi-omics integration, mass spectrometry, networks, Nextflow, nf-core, Python, data science, reproducibility, open science.
He J, Luo S, Xu W, Chen Y, Liu G, Tang J, Yang Y, Zhao B, Ma L, Sheng H, Mao E. Serum proteomic profiling of sepsis patients reveals a protein-based diagnostic model, with metabolomic insights into carbapenem-resistant Klebsiella pneumoniae infection. Front Immunol. 2026;17:1818068. doi:10.3389/fimmu.2026.1818068
Sepsis caused by carbapenem-resistant Klebsiella pneumoniae (CRKP) causes the death of roughly 20β40 % of patients, about twice the rate of a susceptible infection β but blood cultures and susceptibility testing take one to three days, and treatment cannot wait. The study asks whether the patient's own serum molecules can distinguish a resistant from a susceptible infection on day 0.
| Group | n | Description |
|---|---|---|
| Con | 15 | Sepsis, all microbiological cultures negative |
| CSKP | 15 | Sepsis with confirmed carbapenem-susceptible K. pneumoniae (sample IDs KP*) |
| CRKP | 15 | Sepsis with confirmed carbapenem-resistant K. pneumoniae |
The same 45 serum samples were measured on two platforms, which is what makes the integration on Day 3 possible rather than decorative:
| Proteomics | Metabolomics | |
|---|---|---|
| Repository | PXD075261 (ProteomeXchange / iProX) | MTBLS14016 (MetaboLights) |
| Instrument | timsTOF Pro (Bruker) | QTRAP 6500 (SCIEX) |
| Acquisition | diaPASEF, data-independent | MRM, targeted, Β± ionisation |
| Processing | DIA-NN 1.9.2, library-free, MaxLFQ | Vendor MRM integration |
| Features | 1 458 protein groups | 1 073 named metabolites |
| Quality control | 3 pooled injections | 6 pooled injections |
Full documentation of the raw files, the curated tables and their provenance β including two
real data traps the class will meet β is in material/datasets.md.
| Time | Session |
|---|---|
| 9:00β9:30 | Introduction and Housekeeping |
| 9:30β10:00 | From Omics to Multi-omics |
| 10:00β10:30 | β Coffee break |
| 10:30β11:00 | Open Science |
| 11:00β11:30 | Standardising Omics Workflows with Nextflow |
| 11:30β12:30 | π½οΈ Lunch |
| 12:30β13:30 | Introduction to Python |
| 13:30β14:30 | Working with Data in Python |
| 14:30β15:00 | β Coffee break |
| 15:00β16:00 | Visualizing Data in Python |
| Time | Session |
|---|---|
| 9:00β10:00 | Omics: Proteomics and Metabolomics |
| 10:00β10:30 | β Coffee break |
| 10:30β12:00 | Preprocessing Proteomics with quantms/DIA-NN |
| 12:00β13:00 | π½οΈ Lunch |
| 13:00β14:30 | Preprocessing Metabolomics with nf-core/metaboigniter |
| 14:30β16:00 | Proteomics Basic Analysis |
| Time | Session |
|---|---|
| 9:00β10:00 | Metabolomics Basic Analysis |
| 10:00β10:30 | β Coffee break |
| 10:30β11:00 | Multi-omics |
| 11:00β12:00 | Multi-omics I β Integration |
| 12:00β13:00 | π½οΈ Lunch |
| 13:00-13:30 | Introduction to Networks Biology |
| 13:30β14:30 | Multi-omics II β Networks and pathways |
| 14:30β15:00 | β Coffee break |
| 15:00-16:00 | Questions |
Additional material
Visualising Networks β Cytoscape
Networks in Python β Co-abundance Practical
Every hands-on session is a Jupyter notebook that opens in Google Colab with one click on the links above β nothing to install, and the data are downloaded from this repository at run time.
To work locally instead:
git clone https://github.com/Multiomics-Analytics-Group/course_multi-omics_analysis.git
cd course_multi-omics_analysis
pip install -r requirements.txt
jupyter labTwo sessions need more than Python:
- Nextflow pipelines (Day 2 morning) need Java and Nextflow either way. The
proteomics pipeline (
quantmsdiann) additionally needs a container engine (Docker, or Apptainer with a Colab-specific--fakerootfix); the metabolomics pipeline (metaboigniter) uses Conda instead and needs no container engine. Each notebook installs what it needs; seematerial/nextflow_setup.mdfor what to do when a Colab runtime will not cooperate. - Cytoscape (Day 3 morning) is a desktop application β install it beforehand from
cytoscape.org. Instructions:
material/cytoscape.md.
βββ metadata/ clinical and sample metadata for the 45 patients
βββ proteomics/
β βββ data/ protein matrix, annotation, SDRF, published results
β βββ notebooks/ preprocessing (quantms/DIA-NN) and analysis
βββ metabolomics/
β βββ data/ metabolite matrix, annotation, published results
β βββ notebooks/ preprocessing (metaboigniter) and analysis
βββ multiomics/
β βββ data/ published integrated pathway analysis
β βββ notebooks/ integration (SNF, MOFA) and networks
βββ notebooks/ Python, pandas, visualisation and network sessions
βββ slides/ lecture slides
βββ material/ dataset documentation and session instructions
βββ publication/ the paper and its supplementary tables
βββ bin/ scripts that build the curated tables and the notebooks
βββ cheat_sheets/ printable references for Python and its libraries
βββ figures/ logos and images
- He J, et al. Serum proteomic profiling of sepsis patients reveals a protein-based diagnostic model, with metabolomic insights into carbapenem-resistant Klebsiella pneumoniae infection. Front Immunol. 2026;17:1818068. β the course dataset
- Langer BE, et al. Empowering bioinformatics communities with Nextflow and nf-core. Nat Methods. 2025. resource
- Dai C, et al. quantms: a cloud-based pipeline for quantitative proteomics enables the reanalysis of public proteomics data. Nat Methods. 2024;21:1603β1607. resource
- Demichev V, et al. DIA-NN: neural networks and interference correction enable deep proteome coverage in high throughput. Nat Methods. 2020;17:41β44.
- Meier F, et al. diaPASEF: parallel accumulationβserial fragmentation combined with data-independent acquisition. Nat Methods. 2020;17:1229β1236.
- Dai C, et al. A proteomics sample metadata representation for multiomics integration and big data analysis. Nat Commun. 2021;12:5854. β the SDRF standard.
- nf-core/metaboigniter. resource
- Broadhurst D, et al. Guidelines and considerations for the use of system suitability and quality control samples in mass spectrometry assays. Metabolomics. 2018;14:72.
- Dunn WB, et al. Procedures for large-scale metabolic profiling of serum and plasma using gas and liquid chromatography coupled to mass spectrometry. Nat Protoc. 2011;6:1060β1083.
- Wang B, et al. Similarity network fusion for aggregating data types on a genomic scale. Nat Methods. 2014;11:333β337. resource
- Argelaguet R, et al. Multi-Omics Factor Analysis β a framework for unsupervised integration of multi-omics data sets. Mol Syst Biol. 2018;14:e8124. resource
- Cantini L, et al. Benchmarking joint multi-omics dimensionality reduction approaches for the study of cancer. Nat Commun. 2021;12:124.
- BaiΓ£o AR, et al. A technical review of multi-omics data integration methods: from classical statistical to deep generative approaches. Brief Bioinform. 2025;26:bbaf355.
- Shannon P, et al. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res. 2003;13:2498β2504. resource
- Timmons JA, et al. Multiple sources of bias confound functional enrichment analysis of global -omics data. Genome Biol. 2015;16:186.
- Wishart DS, et al. HMDB 5.0: the Human Metabolome Database for 2022. Nucleic Acids Res. 2022;50:D622βD631. resource
- Kanehisa M, et al. KEGG: integrating viruses and cellular organisms. Nucleic Acids Res. 2021;49:D545βD551. resource
- acore β analytical core: filtering, imputation, normalisation, statistics, enrichment and network analysis for omics data
- vuecore β visualisation components
- vuegen β turn a folder of results into a navigable report
Not part of the three-day schedule, but useful for your own projects:
- Basics: Getting started Β· Importing data Β· Jupyter
- Data science: NumPy Β· pandas Β· SciPy Β· scikit-learn
- Visualisation: Matplotlib Β· Plotly Β· Seaborn Β· Bokeh
- learnpython.org β interactive introduction
- Scipy Lectures β Python for scientific computing
- The official tutorial
- Google Colab tutorials β the environment we use
The Python and network notebooks build on material from the Multiomics Analytics Group courses Using Networks to Study Microbes and Omics Data Analysis, and some of them were originally inspired by Python Tsunami at the Center for Health Data Science, University of Copenhagen.
The proteomics and metabolomics analysis sessions build on the Data Science Platform courses dsp_course_proteomics_intro and dsp_course_metabolomics_intro.
We thank He et al. for depositing both omics layers of their cohort publicly.
