Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Today we’re announcing X-Cell — Xaira’s first step toward a virtual cell. 🧬 A foundation model that predicts how gene expression changes under causal perturbations — across cell types, conditions, and even unseen biology. This is not trained on observational atlases. It is trained on interventions. 🧵👇

174,098 Aufrufe • vor 5 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Excited to share our new work. Over the past decade, single-cell genomics has transformed our ability to map cellular systems. But a major question remains: Can we predict how perturbations reshape cellular trajectories over time? In 2018, we first showed that it is possible to predict cellular responses to perturbations — ranging from disease signals to chemical treatments — even in unseen contexts. In 2022, we introduced CPA (MSB 2022; NeurIPS 2022), extending this idea to predict responses to unseen chemical and genetic perturbations, including their combinations. Since then, the field of perturbation modeling has grown enormously. The community has pushed the space forward with many creative ideas and powerful models. It’s exciting to see how fast things are moving — even though many fundamental challenges remain. One of the biggest is that cells are not static. They move through trajectories during development, immune responses, and disease. Yet most current models still predict perturbation effects within a single state, rather than how early perturbations propagate across future states and reshape downstream outcomes. To address this, we developed PerturbGen, a trajectory-aware generative AI model that predicts how genetic perturbations reshape downstream cellular states. Huge credit to the people who made this work possible. Thanks to co-first authors Kevin Ly, Adib Miraki, Tomoya Isobe, AmirHoss3in Vahidi, Delshad Vaghari & Anthony Rostron. Special recognition to Kevin Ly and Adib Miraki for driving this work over the finish line. Grateful for our outstanding collaborators from Haniffa Lab, Bertie Gottgens lab Gosia Trynka and many others — a true cross-institute effort across Cambridge Stem Cell Institute, Open Targets ,Wellcome Sanger Institute and Cambridge University.🎉 PerturbGen learns transcriptional dynamics across cellular trajectories. By introducing perturbations at an early source state, it can simulate how these effects propagate into future states along differentiation trajectories. Scaling this across genes enables the creation of dynamic in silico perturbation atlases — maps of how perturbations reshape biological trajectories over time. We explored this idea across three biological questions. First, in a human in vivo LPS immune challenge, PerturbGen predicted that perturbing a transient IL1B signal dampens downstream inflammatory programs in myeloid cells, with pathway changes reversing signatures observed in an independent IL-1β stimulation experiment. Second, in human hematopoiesis, PerturbGen predicted transcriptional responses to CRISPR transcription factor knockouts and enabled construction of perturbation atlases revealing lineage- and age-specific regulatory programs. These programs could also be linked to human genetics and blood diseases, including recapitulation of signatures associated with ETV6-related thrombocytopenia. Finally, we asked whether perturbation modeling could help improve complex tissue models. We built a dynamic perturbation atlas of human skin organoids to identify perturbations that could guideorganoid cells towardhuman fetal skin states. PerturbGen prioritized activation of Wnt signaling via GSK3β inhibition. Experimental validation confirmed the prediction: treatment with CHIR99021 induced stromal gene programs and shifted organoid fibroblasts toward transcriptional states observed in fetal skin stroma. Together, these results show how trajectory-aware perturbation modeling can connect gene perturbations to developmental programs, human genetics, disease mechanisms, and experimental interventions. More broadly, we think these point toward a future where single-cell atlases become predictive systems. As atlases expand across tissues, developmental windows, and modalities, models like PerturbGen could enable dynamic, virtual perturbation atlases— allowing us to simulate interventions, generate hypotheses, and design experiments before stepping into the lab. Preprint Code Excited to see how the community builds on this work.

Mo Lotfollahi

17,107 Aufrufe • vor 5 Monaten

How does an embryo reliably "compute" its form - "cell by cell" - using only local interactions and mechanics, yet produce a precise global body plan? I’m excited to share our Nature Methods paper "MultiCell: geometric learning in multicellular development", presenting #AIxBiology research led by Haiqian Yang and the result of a great collaboration with Ming Guo, George Roy, Tomer Stern, Anh Nguyen and Dapeng Bi. A long-standing challenge in developmental biology is to predict how thousands of cells collectively self-organize as tissues fold, divide, and rearrange. In MultiCell, we represent a developing embryo as a dual graph that unifies two complementary views of tissue mechanics with single-cell resolution: cells as moving points (granular) and cells as a connected foam (junction network). This lets the model learn dynamics from both geometry and cell–cell connectivity. On whole-embryo 4D light-sheet movies of Drosophila gastrulation (~5,000 cells), our model predicts key cell behaviors and the timing of events, including junction loss, rearrangements, and divisions with high accuracy, at single-cell resolution. Beyond prediction, the same representation supports robust time alignment across embryos and offers interpretable activation maps that highlight the morphogenetic "drivers" of development. The broader goal is a foundation for cell-by-cell forecasting in more complex tissues, and eventually for detecting subtle dynamical signatures of disease. Kudos to the team for this inspiring collaboration with brilliant researchers to push the boundary of AI for biology! Citation: Yang, H., Roy, G., Nguyen, A.Q., Buehler, M.J., et al. MultiCell: geometric learning in multicellular development. Nature Methods (2025), DOI: 10.1038/s41592-025-02983-x Code/data links are in the manuscript.

Markus J. Buehler

388,166 Aufrufe • vor 7 Monaten

🔬 Exciting News! Our manuscript, "scGPT: toward building a foundation model for single-cell multi-omics using generative AI" is now finally published in Nature Methods (Nature Methods) 🎉 !!! (Re-)Introducing scGPT: A transformative foundation model engineered for single-cell omics analysis. Developed through the analysis of over 33 million human cells, scGPT sets a new benchmark for application versatility, offering both fine-tuning and zero-shot capabilities. Since its preprint in May 2023, scGPT has significantly impacted the field, evidenced by 13K+ installations, 600+ GitHub stars 🌟, and 40+ citations before its official publication! scGPT has been validated by numerous benchmark studies as a leading foundation model in single-cell analysis. Its pre-trained embeddings extend its utility beyond single-cell studies, enhancing a variety of downstream tasks including protein enrichment and genetic perturbation predictions. Some key updates lately: ---Expanded zero-shot applications for efficient reference mapping and integration, now with CellXGene census integration. ---Advanced perturbation analysis capabilities, including genome-scale perturb-seq data analysis and bulk sequencing data generalization. ---Upgraded scGPT package, offering versatile model loading compatible with PyTorch and flash-attn, for both GPU and CPU. ---Cloud-based scGPT applications for reference mapping, cell annotation, and gene regulatory network inference are available on ---Integration with Hugging Face for easier model training. Limitations: scGPT is an early foray into foundation models for single-cell omics, facing challenges like limited zero-shot learning in some tasks, pretraining constraints, data quality issues, and evaluation limitations. See our Supplementary Notes for details. 🚀 Future Work? Short-Term Goals: 1. Releasing a Mouse Model for broader analysis. 2. Developing a comprehensive evaluation suite for foundation models in single-cell analysis. 3. Creating a foundation model for single-cell spatial omics. 4. Enhancing zero-shot capacity by integrating scGPT with RAG (e.g., knowledge graphs). Long-Term Goals: 1. Expanding scGPT for comprehensive single-cell multi-omics analysis. 2. Developing an in-silico perturbation model for predicting genetic perturbation effects. 3. Merging scGPT with multi-modal genomic sequence models for a deeper understanding of cell biology. 📚 Access the paper on Nature Methods: 🔬Preprint in Bioarixv: 💻 All our codes/data/weights are open source: Wholehearted congratulations to all the authors, especially the two co-first authors, Haotian (Haotian Cui ) and Chloe (ChloeXWang), who are really the emerging superstars in AI and biology! Vector Institute Peter Munk Cardiac Centre AI U of T Department of Computer Science Department of Laboratory Medicine & Pathobiology University Health Network University of Toronto #scGPT #GenerativeAI #AI4Science #Combio #opensource

Bo Wang

199,725 Aufrufe • vor 2 Jahren

🚀 Introducing scGPT-spatial! 🧬🌍 A game-changing spatial-omic foundation model, built on the powerful scGPT framework with MoE (mixture of experts) and continually pretrained on a massive 30 million spatial single-cell profiles! 🧠 What’s the challenge? Spatial transcriptomics is next-level complex—not only must we model single-cell/spot profiles, but we also need to capture intricate spatial relationships while handling diverse sequencing protocols (imaging-based vs. sequencing-based). 🔥 Why scGPT-spatial? ✨ A Spatial-omic Foundation Model with Continual Pretraining – Built on scGPT’s robust initialization, it unlocks spatial context in tissues. ✨ SpatialHuman30M Dataset – The largest curated dataset: 30M profiles from Visium, Visium HD, Xenium, and MERFISH across 821 slides. ✨ Revolutionary MoE Decoders – A cutting-edge Mixture of Experts (MoE) architecture for protocol-aware gene expression decoding. ✨ Spatially-Aware Training Strategy – A neighborhood-based masked reconstruction approach to capture complex cell-type colocalization. ✨ Multi-Modal & Multi-Slide Integration – Seamless clustering & spatial domain identification across slides and modalities. ✨ Cell-Type Deconvolution & Gene Imputation – Unlocks cross-resolution & cross-modality harmonization with fine-tuned embeddings. 📄 Read the preprint: 💻 Explore the code/weights: #SpatialTranscriptomics #SingleCell #AIResearch #MachineLearning #SpatialData Huge shoutout to the incredible PHD students Chloe (ChloeXWang) and Haotian (Haotian Cui) for leading this groundbreaking project! 🎉 Massive thanks to our amazing co-authors Andrew, Ronald, and Hani (Hani Goodarzi) from Arc Institute—this work wouldn't have been possible without you! 👏

Bo Wang

59,066 Aufrufe • vor 1 Jahr

This is crazy! This is some of the complex systems working inside your cells. God's Design inside every single cell in your body. Cells cannot arise through evolution. The minimum viable cell requires: - Around ~500,000 lines of coded information (DNA) - Close to ~500 unique protein products (the building blocks of all those cellular machines and systems) - Total of ~20k-50k total proteins all working perfectly together - Around ~30-40 regulatory systems guiding all those interactions The cell requires all those parts & systems, or it doesn't function. If we do the math, there are about ~10^70,000 possible interactions in this cell. Interactions are things like: - energy production - waste removal - protein creation - system repair The odds are incomprehensible. Evolutionists will argue the odds are misleading, because it evolves gradually via step-by-step trial & error. But the cell REQUIRES a minimum set of parts & systems to function - without all of these in place, together, from the beginning, it dies. Therefore, the cell cannot evolve through step-by-step evolutionary processes, because there is no reproduction + mutation to drive evolutionary change until the cell is complete. Cells are a massive problem for Evolutionism, because they are such obvious signs of Intelligent Design. Even the famous atheist Richard Dawkins admitted that, "Biology is the study of complicated things that have the appearance of having been designed with a purpose." Biology seems designed because it IS designed. God's Divine Design becomes more obvious, the closer we look.

Divinely Designed

32,299 Aufrufe • vor 5 Monaten

Single-cell technologies now let us profile entire transcriptomes in individual cells. But how do we make sense of this complexity in a biologically meaningful way? Many methods summarise cells into a single embedding, but this often comes at the cost of interpretability, especially when multiple gene programs are active at once. We developed Tripso, a self-supervised transformer model that represents cells through multiple gene program-specific embeddings, while also uncovering new programs directly from the data. Instead of collapsing biology into a single vector, Tripso decomposes cell state into multiple representations, each reflecting a different gene program. We explored this across multiple systems. In human hematopoiesis, spanning development to aging, Tripso identified distinct age-associated program activity, including stronger JAK-STAT signalling in early life and dynamic IKZF1-related changes during B cell maturation. By comparing in vitro culture conditions with in vivo hematopoietic stem cell states, Tripso suggested that targeting the SEC61 translocon could enhance stem cell maintenance ex vivo, a prediction that we subsequently validated experimentally. In parallel, we identified a previously uncharacterised tissue-resident memory T-cell program associated with atopic dermatitis and mapped it to distinct spatial immune niches Together, these results show how modelling cells through gene programs can lead to interpretable and experimentally testable insights. More broadly, this work points toward a more interpretable and biologically grounded models of cell state. As single-cell datasets continue to grow, we hope approaches like Tripso will help bridge the gap between data-driven representations and biological insight. This work wouldn’t have been possible without the contributions of an amazing team. Thank you to co-first authors Marie, Tomoya Isobe, Amirhosein Vahidi, Carlo Leonardi, and everyone from roser's Lab, Haniffa Lab, Nicola Wilson and Bertie Gottgens's Lab, bringing together expertise across Cambridge Stem Cell Institute, Open Targets, Wellcome Sanger Institute and Cambridge University. Marie is one of the very best PhD students I have ever supervised. She is truly a force of nature, exceptionally resourceful, deeply innovative, and one of the most impressive scientists I have worked with. I am immensely proud of her and all that she has accomplished. As she begins her internship at Genentech , I have no doubt she will do amazing work there and continue to make her mark. paper: code:

Mo Lotfollahi

22,498 Aufrufe • vor 4 Monaten