Le Cao lab’s cover photo
Le Cao lab

Le Cao lab

Biotechnology Research

Parkville, Victoria 280 followers

Computational statistics and biology

About us

Welcome to the research lab of Prof. Kim-Anh Lê Cao. The Le Cao Lab, established in October 2008, is a globally recognised research group based at Melbourne Integrative Genomics (MIG) and the School of Mathematics and Statistics at the University of Melbourne. Led by Professor Kim-Anh Lê Cao, recipient of three consecutive NHMRC fellowships since 2014 and the Australian Academy of Science Moran Medal, our team innovates at the forefront of computational statistics and biological data integration. Our core mission is to develop efficient computational solutions to interpret large-scale ‘omics datasets. We are the creators of mixOmics, a leading R toolkit for multivariate data analysis. Ranked in the top 5% of Bioconductor packages with over 30,000 users yearly, our tools are considered the gold standard for data integration. Our research impacts are diverse and tangible. Our algorithms have been applied to safeguard the Great Barrier Reef and improve productivity in the dairy industry, as well as to accelerate medical breakthroughs, ncluding identifying biomarkers for Alzheimer's disease, cancer, and diabetes. We empower consortia and industry partners globally to translate complex data into real-world decisions. Beyond methodological development, we are a hub for collaboration and training. We have secured over AUD $73M in competitive research funding, including the ARC Centre of Excellence in Quantum Biotechnology. We are deeply committed to mentoring the next generation of data analysts, with alumni successfully establishing careers in international research institutions.

Website
https://lecao-lab.science.unimelb.edu.au/
Industry
Biotechnology Research
Company size
11-50 employees
Headquarters
Parkville, Victoria
Type
Nonprofit
Founded
2008
Specialties
Multi-Omics Data Analysis, Bioinformatics, Data Integration, Multivariate Statistics, Biotechnology, Data Analysis Platform, Genomics, Transcriptomics, Proteomics, Metabolomics, Dimensionality Reduction, Life Sciences, Pharmaceutical Research, Clinical Research, Data Visualisation, Agricultural Research, Scientific Software, and mixOmics

Locations

  • Primary

    The University of Melbourne, Grattan Street

    Parkville, Victoria 3010, AU

    Get directions

Employees at Le Cao lab

Updates

  • A great milestone for the lab! Marko Terzin has been visiting us during his PhD to learn about our mixOmics tools and is now a postdoc in the lab. We are looking forward to even more impactful results from this collaboration!

    🌊 Major milestone for ocean science! 🌊 Our team’s study comprehensively mapping the planktonic microbiome of the Great Barrier Reef is published today in Nature, and featured in The Guardian! 🗞️✨ To protect an ecosystem as complex as the Great Barrier Reef, we first have to understand the microscopic life driving its health. By analyzing seawater across 48 reefs, our multi-institutional team built the Great Barrier Reef Microbial Genomes Database (GBR-MGD): 🦠 580+ bacterial & archaeal species new to science 🧬 360,000+ distinct marine viruses 🌿 Chromosome-level genomes of key reef microalgae 🛡️ A comprehensive baseline to track reef health and predict responses to environmental stress 📊 Cutting through the data noise with mixOmics MINT method Analyzing multi-omic data sampled across thousands of kilometers and multiple seasons generates an overwhelming amount of information. Our biggest analytical hurdle was to separate true, underlying biological patterns from spatial and temporal noise. On the analytical front, Marko Terzin and our team at Le Cao lab (Faculty of Science, Melbourne Integrative Genomics) leveraged our mixOmics MINT method to integrate complex multi-omic datasets while explicitly controlling for geographical and seasonal variability. This statistical approach allowed us to isolate robust microbial signatures that reflect ecosystem dynamics across the entire reef system, providing a reliable framework for future monitoring. A massive effort by an extraordinary team across The University of Queensland Australian Institute of Marine Science, University of Melbourne, James Cook University, and University of Tasmania! 🥂 📖 Read the Nature paper: https://lnkd.in/gecgH6kF 📰 Read The Guardian feature: https://lnkd.in/gmDxD2_Q Co-authors: Steven J. Robbins, PhD | Marko Terzin | Katherine Dougan | Julian Zaugg | Sara Bell | Patrick Laffy | Pam Engelberts | Kim-Anh Lê Cao | Renee Gruber | @Nicole S. Webster | David Bourne | Philip Hugenholtz | Yun Kit YEOH

    • Credit: DETSI QLD gov
  • Le Cao lab reposted this

    A new collaboration between our lab (analysis led by Jiadong Mao) and Sarah-Maria Fendt's lab (work led by Xiaozheng L.), facilitated by Mark Dawson. We are gearing up to new projects in that space (so to speak!).

    Starving metastatic cancer cells! A great example of spatial multiomics and spatial medicine – excited to share our collaborative work with Prof Sarah-Maria Fendt's lab at VIB–KU Leuven Center for Cancer Biology, now published in Cancer Discovery. The study reveals how cancer cells hijack healthy lung cells, specifically alveolar type II (AT2) cells, AT2 are reprogramed to ramp up lipid production that fuels metastatic growth. Crucially, blocking this lipid supply reduced lung metastasis, pointing to a strategy of targeting the tumour microenvironment rather than cancer cells directly. PhiSpace was used to validate one of the central claims of the paper using a previously under-utilised spatial transcriptomics sample. we mapped how AT2 cells shift their molecular states in the presence of cancer. Despite the relatively low resolution of the data (multicellular Visium), we managed to extract some strong signals. This great example reflects the direction I'm most passionate about: working with domain experts to crack challenging biological problems, from a computational angle. Congratulations to first author Xiaozheng L., senior author Sarah-Maria Fendt and the whole team! TDLR version: https://lnkd.in/gBgWBq5G Full text: https://lnkd.in/gh25RmYz

    • No alternative text description for this image
    • Graphical abstract
    • Validation using spatial transcriptomics with PhiSpace
  • Our lab is hiring! We are very excited to welcome new members to the team, to work at the cutting-edge of omics data analytics, and at the frontier of quantum computing!

    Hiring: Data Analyst & Research Fellow, Integrative and Quantum Omics (Melbourne, Australia) How do we unravel the complexity of biological systems? In the Le Cao lab, we believe the answer lies at the intersection of advanced statistics, multi-omics integration, and the emerging frontier of quantum computing. We are looking for two driven individuals to join our team based at Melbourne Integrative Genomics (MIG) to push the boundaries of computational biology. Both positions are full-time, fixed-term roles for 2 years. 🔬 The roles - Data Analyst (Omics): You will conduct end-to-end analytics on diverse datasets (multi-omics, microbiome, bulk and single cell). - Postdoctoral Research Fellow (Quantum Omics): Working within the ARC Centre of Excellence in Quantum Biotechnology (QUBIC), you will spearhead method development at the exciting intersection of quantum computing and integrative omics. 🎯 Who this suits - For the Analyst: A detail-oriented expert in R and Python with a Master’s degree in Bioinformatics or Statistics and a passion for reproducible research - For the Fellow: A highly motivated PhD with expertise in both omics data analysis and quantum computing methods who wants to lead innovative research. - Both roles: Collaborative spirits with an active GitHub profile showcasing their code. - Visa sponsorship is available for the Research Fellow position only; current valid Australian working rights are required for the Data Analyst role. 🌍 What you gain - Competitive Compensation plus 17% superannuation. - World-Class Environment: Based in the Parkville precinct, the highest concentration of biological and medical research in Australia. - Professional development opportunities, and access to a vibrant community of 50+ researchers at MIG. If you are ready to unlock your career potential in a world-class team, apply via the links below! 👇 📅 Applications close  6 April 2026. 11:55 PM; Melbourne time zone. - Data analyst: https://lnkd.in/gCUSqzWm - Research Fellow: https://lnkd.in/g6Wmhbgc 📧 Email me if you want to check the suitability of the role. Please share widely through your network!

    • No alternative text description for this image
  • Exciting to see this work from Jiadong Mao and Yinuo Sun as a preprint! While most integration tools look for the "lowest common denominator" across data (like mixOmics methods), DIVAS performs a comprehensive hierarchical search to find signals that might only exist in a subset of modalities. In our COVID-19 case study, this approach identified specific immune and metabolic dysregulation patterns, like the interconnected network of chemokines and metabolites driving the "cytokine storm", that conventional methods would have simply overlooked. It’s a powerful reminder that the most interesting biology often happens between the layers. Huge thanks to Jiadong Mao and Yinuo Sun for making this robust computational framework available as an open-source R package for the wider research community.

    Check out our new software paper on integrating >= 3 omics data using #DIVAS 😃 When integrating >= 3 omics modalities, a common challenge is: there are plenty of methods that tell you what are the modes of variation shared by all of them; there are some methods that tell you what are the modes of variation individual to each of them; but what about the modes of variation that are only partially shared? 🧐 Partially shared variations are important because not all omics are qually informative. You might, for example, expect more biological variations shared between RNAs and proteins, but comparatively less shared between proteins and metabolomics. If you use conventional methods you might miss out those partially shared variations. DIVAS uses a matrix factorisation approach that does a comprehensive search for significant modes of variations among all combinations of omics types: jointly shared, partially shared and individual. Its data-driven algorithm requires very minimal manual tuning. You don't need to specify any number of components! 💪 The DIVAS R package project is led by Yinuo Sun, with strong support from the original developers J. S. Marron and his team at UNC Chapel Hill. Corresponding authors Kim-Anh Lê Cao & Jiadong Mao https://lnkd.in/ghCaBfVu #DataIntegration #multiomics #DataFussion #Bioinformatics #Statistics

    • No alternative text description for this image
    • No alternative text description for this image
  • View organization page for Le Cao lab

    280 followers

    Seeing our research and development on mixOmics recognised by the Faculty of Science with the inaugural Impact Award is a fantastic milestone to finish 2025! This recognition belongs to the entire team, both past and present. While many of our contributors have since moved on to new chapters, their impact remains at the core of this success. We are incredibly grateful to Dr Florian Rohart GAICD, @benoit gautier, Amrit Singh, Eva Hamrud, Al J Abadi, @Max bladen. Looking forward to carrying this momentum into 2026! Melbourne Integrative Genomics | Faculty of Science | Moira O'Bryan

    • No alternative text description for this image
  • View organization page for Le Cao lab

    280 followers

    Great opportunity at the University of Melbourne to work with cutting edge metabolomics.

    💡 Job Opportunity Alert!! 💡 The Metabolomics Australia node at the University of Melbourne is looking a talented individual with a passion for metabolomics to join our team. This bioinformatics role will focus on building robust pipelines, analyzing and interpreting high-dimensional metabolomics data, and collaborating closely with other team members and researchers to generate amazing data. Key skills: • Metabolomics data analysis (LC–MS/GC–MS) • Bioinformatics/statistical workflows (R, Python, or similar) • Experience with multi-omics data is a plus If you’re excited about applying computational methods to cutting-edge biological questions, we’d love to hear from you. We particularly encourage individuals with strong mass spectrometry experience, who have developed computational skills along the way to apply. Please note that we are unable to provide visa sponsorship, so candidates will have to already possess Australian working rights. #metabolomics #bioinformatics #massspec https://lnkd.in/gc6wkQYg

  • View organization page for Le Cao lab

    280 followers

    Unsupervised checks are done. Now, we move to supervised validation. As mixOmics instructors, we always remind our learners that checking for structure is a prerequisite, but it does not guarantee a discriminant signal (using PCA, see Post 3, https://lnkd.in/gVzrtdtm). That is where PLS-DA comes in. We must assess each layer individually first. Using PLS-DA and its sparse variant (sPLS-DA) allows us to quantify the "baseline" power of each omics layer and identify potential discriminant features early on. If the single-omics analysis fails to separate your groups, integration is unlikely to fix it. Follow the guide below 👇 https://lnkd.in/gHj5MvjW ——— mixOmics Pro  |  mixOmics  |  Le Cao lab | Kim-Anh Lê Cao #bioinformatics #computationalBiology #dataScience #RStats #scienceCommunication ———

    View organization page for mixOmics Pro

    742 followers

    ⭐  Single-omics first — integration must ADD value, not dilute it  ⭐ More omics data is not automatically better. Before combining layers, you need to know which ones actually carry a discriminative signal. We use PLS-Discriminant Analysis (PLS-DA) to classify samples into predefined groups and to evaluate cross-validated discriminative performance for each omics layer on its own. 1️⃣  ESTABLISH A BASELINE   •  run plsda() on each dataset individually   •  use perf() to assess cross-validated classification error   •  a strong layer (e.g. low error / ~80% accuracy) becomes the baseline   •  any integrated model must beat this baseline OR deliver additional, interpretable biology to be worth the effort 2️⃣  IDENTIFY WEAK LAYERS - noise dilutes the signal   •  if a layer shows poor class separation (discrimination) and high error, treat it as weak   •  decide: does this layer add interpretable biological structure, or just noise?   •  only integrate weak layers if they clearly strengthen the overall model - otherwise they can sink your integration performance 🏃🏽  ACTION STEPS (using mixOmics) • plsda() on each dataset individually • perf() to assess cross-validated error • compare baselines before deciding which layers to integrate 🔑  GOLDEN RULE More omics ≠ better. Integration is an investment – only integrate layers that strengthen the signal and the biology ✒️  COMMENT Do you always check your single-omics baselines first, or jump straight to the "whole"? ♥️  Like and follow to learn more 🔥  Get multi-omics tips in your inbox: https://mixOmics.pro ——— 💫 Part 4/6 of our "Building a Robust Omics Study" series. 📢 Next week in Part 5: True integration vs. simple correlation. mixOmics Pro | mixOmics | Le Cao lab | Kim-Anh Lê Cao | Mike Rennie ——— #multiOmics #omics #omicsIntegration #bioinformatics #computationalBiology #dataScience #lifeSciences #biotech #systemsBiology #metabolomics #proteomics #transcriptomics #microbiome #PhDLife #scienceCommunication #RStats #LeCaoLab #mixOmics #mixOmicsPro ———

    • No alternative text description for this image
  • View organization page for Le Cao lab

    280 followers

    PCA is still the simplest and most revealing first step in multi-omics quality control. In the lab, we always run PCA per omics layer before combining datasets. Refer to this great visual guide from mixOmics Pro on what to look for: outliers, batch drift, and low-information layers. ——— mixOmics Pro  |  mixOmics |  Kim-Anh Lê Cao #bioinformatics #computationalBiology #dataScience #RStats #scienceCommunication ———

    View organization page for mixOmics Pro

    742 followers

    ⚠️  STOP ⚠️ PCA first ➔ because integration amplifies errors‼️ 🚨 Run PCA on each omics layer BEFORE combining datasets 1️⃣  OUTLIERS   •  Drifting samples = either biology or technical failure   •  Keep biological outliers, remove technical ones 2️⃣  A BATCH EFFECT   •  If samples cluster by date/lab/technician instead of treatment, your “signal” may be a workflow artifact   •  Correct batch effects before integrating 3️⃣  RELEVANCE   •  A layer with no structure only dilutes the integrated model   •  If it’s noise, exclude it 🏃🏽 ACTION STEPS (using mixOmics):   •  Run PCA (unsupervised) to reveal raw structure   •  Colour by treatment to check biological grouping   •  Colour by batch/date to expose technical artefacts 🔑 GOLDEN RULE  Integrate only layers that strengthen the signal — otherwise, integration makes things worse ✒️ COMMENT What’s the strangest batch effect you’ve seen? 🔥 UPDATES Get multi-omics tips in your inbox: https://mixOmics.pro — 💫 Part 3/6 of our "Building a Robust Omics Study" series 📢 Next week: Why single-omics analysis always comes first mixOmics Pro | mixOmics  |  Le Cao lab  |  Kim-Anh Lê Cao | Mike Rennie#multiOmics #Omics #omicsIntegration #Bioinformatics #computationalBiology #dataScience #LifeSciences #Biotech #systemsBiology #metabolomics #proteomics #transcriptomics #microbiome #PhDLife #scienceCommunication #RStats #LeCaoLab #mixOmics #mixOmicsPro

    • No alternative text description for this image
  • Le Cao lab reposted this

    🎉Congratulations to Yeganeh Khazaei! Our latest manuscript, led by Yeganeh Khazaei, is now live in BMJ Open! “Development and validation of diagnostic and prognostic prediction tools for dental caries in young children through prospective and cross-sectional observational studies: a protocol” This new publication from Melbourne Integrative Genomics and the Melbourne Dental School at the University of Melbourne outlines our approach to developing and validating prediction models for dental caries (tooth decay) in young children—integrating environmental, physical, behavioural, and biological early-life data to identify at-risk children and inform early prevention strategies. Authors: Yeganeh Khazaei, Saritha Kodikara, Catherine Butler, Nicole Messina, Kim-Anh Lê Cao, Stuart Dashper and Mihiri Silva Read the manuscript on BMJ Open: https://lnkd.in/egC7MS_6 #DentalCaries #EarlyChildhoodHealth #MicrobiomeResearch #ChildOralHealth

    • No alternative text description for this image

Similar pages