Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

We NOETIK built a new foundation model for spatial biology: OCTO-VirtualCell. This model is prompted to simulate single cell gene expression in "virtual cells," which are placed within real, intact patient tissue. Read about & explore these simulations below:

39,840 Aufrufe • vor 1 Jahr •via X (Twitter)

10 Kommentare

Profilbild von Daniel Bear
Daniel Bearvor 1 Jahr

You can get an intuitive feel for how virtual cell simulations work through Celleporter, an interface for viewing OCTO-vc predictions in patient samples: Or read how we're using this model for drug discovery in a technical report:

Profilbild von Daniel Bear
Daniel Bearvor 1 Jahr

OCTO-vc was trained on a @NOETIK_ai-generated spatial transcriptomics dataset of 40M cells across more than 1000 tissue samples from patients with cancer. More about how we generate massive, high-quality datasets for self-supervised learning:

Profilbild von Daniel Bear
Daniel Bearvor 1 Jahr

OCTO-vc simulates based on a prompt, which consists of 1. a few genes set as on or off in the virtual cell 2. the local and patient-level biological context, such as genes expressed by nearby cells. The model then predicts expression of the remaining genes in the virtual cell.

Profilbild von Daniel Bear
Daniel Bearvor 1 Jahr

Virtual cell simulations reveal a huge amount of biology. One cool example: we can use virtual B cells to see the spatial and functional structure of tertiary lymphoid structures, sites where B cells mature and produce antibodies in response to infection or cancer.

Profilbild von Daniel Bear
Daniel Bearvor 1 Jahr

In another case, OCTO-vc predicts that virtual killer (CD8) T cells will be in different functional states in different regions of a patient's tumor. This heterogeneity is central to @NOETIK_ai 's mission of finding cancer immunotherapies that work in each patient.

Profilbild von Daniel Bear
Daniel Bearvor 1 Jahr

Counterfactual simulations with virtual cells can reveal even more differences between patients. Most patients' killer T cells are predicted to be more effective the more cancer cells are presenting antigens on their surface -- the expected biology. But some patients don't show this response. This could be a clue as to why many patients don't respond to existing immunotherapies.

Profilbild von Daniel Bear
Daniel Bearvor 1 Jahr

We can scale up this approach to run a "virtual screen" on an entire patient cohort. In this one, we first found that virtual killer T cells are predicted to be in a less functional state in a specific set of patients known to be unresponsive to immunotherapy.

Profilbild von Daniel Bear
Daniel Bearvor 1 Jahr

Then, we can simulate the knockdown of different genes, one by one, in each patient's tissue. Some of these simulated interventions partially "rescue" the virtual T cells, as OCTO-vc predicts they'll become more effective at killing cancer cells. These are ingredients for finding the right drugs for each patient.

Profilbild von Daniel Bear
Daniel Bearvor 1 Jahr

This is a transformative moment for understanding the world of biology and disease. So much of biology is spatial: how cells come together to form functional tissues and how they communicate with each other. But our brains didn't evolve to make sense of thousands of genes and proteins expressed across labyrinthine structures -- and there's no army of data labelers coming to tell us what it all means. Self-supervised learning on real patient data is the way. We've now crossed a threshold where data-generating technology and machine learning methods, combined, are powerful enough to reveal biology deeper than we've ever been able to see before. We're beyond excited to explore this new world and we're looking for collaborators!

Profilbild von Daniel Bear
Daniel Bearvor 1 Jahr

It's awesome that just as we've started building these models, people are getting excited about using ML-based simulation to discover new biology -- with the idea of a "virtual cell" as a unifying framework: We're tackling a particular problem here -- focused especially on learning spatial and functional relationships between *different* cells -- where the concept of a "virtual cell" as a unit of simulation makes complete sense. We're excited about bringing in other data and data modalities and simulating other aspects of cell + patient biology, and looking forward to seeing what problems others are using the virtual cell concept to tackle.

Ähnliche Videos

🚀 Introducing scGPT-spatial! 🧬🌍 A game-changing spatial-omic foundation model, built on the powerful scGPT framework with MoE (mixture of experts) and continually pretrained on a massive 30 million spatial single-cell profiles! 🧠 What’s the challenge? Spatial transcriptomics is next-level complex—not only must we model single-cell/spot profiles, but we also need to capture intricate spatial relationships while handling diverse sequencing protocols (imaging-based vs. sequencing-based). 🔥 Why scGPT-spatial? ✨ A Spatial-omic Foundation Model with Continual Pretraining – Built on scGPT’s robust initialization, it unlocks spatial context in tissues. ✨ SpatialHuman30M Dataset – The largest curated dataset: 30M profiles from Visium, Visium HD, Xenium, and MERFISH across 821 slides. ✨ Revolutionary MoE Decoders – A cutting-edge Mixture of Experts (MoE) architecture for protocol-aware gene expression decoding. ✨ Spatially-Aware Training Strategy – A neighborhood-based masked reconstruction approach to capture complex cell-type colocalization. ✨ Multi-Modal & Multi-Slide Integration – Seamless clustering & spatial domain identification across slides and modalities. ✨ Cell-Type Deconvolution & Gene Imputation – Unlocks cross-resolution & cross-modality harmonization with fine-tuned embeddings. 📄 Read the preprint: 💻 Explore the code/weights: #SpatialTranscriptomics #SingleCell #AIResearch #MachineLearning #SpatialData Huge shoutout to the incredible PHD students Chloe (ChloeXWang) and Haotian (Haotian Cui) for leading this groundbreaking project! 🎉 Massive thanks to our amazing co-authors Andrew, Ronald, and Hani (Hani Goodarzi) from Arc Institute—this work wouldn't have been possible without you! 👏

Bo Wang

59,008 Aufrufe • vor 1 Jahr

Excited to share TERRA, a tissue world model 🧬 Over ~1.5 years we ran a large data-generation + modelling effort to build a world model for human tissues, pretrained on 112M cells from spatial transcriptomics (mostly Xenium 5000-plex + public data). It's built on one of the largest human spatial transcriptomics corpora assembled to date, spanning 20 tissues across development, health and 26 disease conditions, ~two-thirds newly generated in-house. Why a "world model" for tissue? Images have universal representations (ViT/DINOv3), so do proteins (ESM, Alex Rives) and pathology (UNI, Faisal Mahmood). We've worked hard to build something similar for human tissue: one model that captures its multi-scale logic, genes → cells → their native microenvironments. Like the JEPA approach Yann LeCun has championed, TERRA learns by prediction in embedding space, but for human tissue. How it works: it tokenises each cell together with its nearest neighbours into one sequence while keeping every gene's identity, then masks part of a neighbourhood and predicts the representation of the hidden part, not raw noisy counts. From one backbone it reads out three scales, gene embeddings (what a gene is doing in a cell and its niche), cell embeddings (cell type and state) and neighbourhood embeddings (the niche), and because it keeps gene-level resolution it can knock a gene out in silico and predict the response. Applied entirely zero-shot, TERRA maps and perturbs human tissue across unseen organs, diseases and technologies, outperforming existing spatial approaches. Three take-homes: 1️⃣ One model, any tissue. A single pretrained backbone provides tissue representations zero-shot, handling genes, cells and niches across organs and platforms, off the shelf. 2️⃣ New biology, development to clinic. We built a new spatial atlas of the developing human pancreas and found an islet-associated capillary state that looks like a precursor of mature islet vasculature. In kidney, TERRA's in silico knockouts predicted the tissue-injury programme from cancer immunotherapy (checkpoint blockade), confirmed in treated kidneys, detected in blood, and linked to declining kidney function. 3️⃣ A grammar of tissue architecture. By coupling each cell's state to its niche, TERRA defines recurring cross-organ "archetypes" of macrophage neighbourhoods, including a tumour-boundary niche that tracks poor survival in kidney cancer. TERRA is already in use: it powered our recent skin atlas of hidden immune-memory niches ( with more studies coming soon. This was an amazing collaboration between clinicians, machine-learning scientists and cell biologists 🙏 Led by Sebastian Birk, Vali Sanian Amirhosein Vahidi, Samuel Ogden, Daniyal Jafree, Adib Miraki, Carlo Leonardi and Arpit Merchant, with Lassi Paavolainen, Menna Clatworthy, Omer Ali Bayraktar, Muzlifah Haniffa, Tom Mitchell and Mostafa Bakhti. Huge thanks too to everyone who shared data and helped along the way. What excites me most is seeing how the community builds on this. The model, code and tutorials are all public, so anyone can run TERRA on their own tissues, extend it, or build new models on top. Huge thanks to the whole team across Wellcome Sanger Institute and our many collaborators. 📄 Paper: 💻 Code: 🤗 Model: #SpatialTranscriptomics #SpatialGenomics #FoundationModels #AI4Science #MachineLearning #ComputationalBiology #SingleCell #WorldModels

Mo Lotfollahi

35,293 Aufrufe • vor 2 Tagen

🔬 Exciting News! Our manuscript, "scGPT: toward building a foundation model for single-cell multi-omics using generative AI" is now finally published in Nature Methods (Nature Methods) 🎉 !!! (Re-)Introducing scGPT: A transformative foundation model engineered for single-cell omics analysis. Developed through the analysis of over 33 million human cells, scGPT sets a new benchmark for application versatility, offering both fine-tuning and zero-shot capabilities. Since its preprint in May 2023, scGPT has significantly impacted the field, evidenced by 13K+ installations, 600+ GitHub stars 🌟, and 40+ citations before its official publication! scGPT has been validated by numerous benchmark studies as a leading foundation model in single-cell analysis. Its pre-trained embeddings extend its utility beyond single-cell studies, enhancing a variety of downstream tasks including protein enrichment and genetic perturbation predictions. Some key updates lately: ---Expanded zero-shot applications for efficient reference mapping and integration, now with CellXGene census integration. ---Advanced perturbation analysis capabilities, including genome-scale perturb-seq data analysis and bulk sequencing data generalization. ---Upgraded scGPT package, offering versatile model loading compatible with PyTorch and flash-attn, for both GPU and CPU. ---Cloud-based scGPT applications for reference mapping, cell annotation, and gene regulatory network inference are available on ---Integration with Hugging Face for easier model training. Limitations: scGPT is an early foray into foundation models for single-cell omics, facing challenges like limited zero-shot learning in some tasks, pretraining constraints, data quality issues, and evaluation limitations. See our Supplementary Notes for details. 🚀 Future Work? Short-Term Goals: 1. Releasing a Mouse Model for broader analysis. 2. Developing a comprehensive evaluation suite for foundation models in single-cell analysis. 3. Creating a foundation model for single-cell spatial omics. 4. Enhancing zero-shot capacity by integrating scGPT with RAG (e.g., knowledge graphs). Long-Term Goals: 1. Expanding scGPT for comprehensive single-cell multi-omics analysis. 2. Developing an in-silico perturbation model for predicting genetic perturbation effects. 3. Merging scGPT with multi-modal genomic sequence models for a deeper understanding of cell biology. 📚 Access the paper on Nature Methods: 🔬Preprint in Bioarixv: 💻 All our codes/data/weights are open source: Wholehearted congratulations to all the authors, especially the two co-first authors, Haotian (Haotian Cui ) and Chloe (ChloeXWang), who are really the emerging superstars in AI and biology! Vector Institute Peter Munk Cardiac Centre AI U of T Department of Computer Science Department of Laboratory Medicine & Pathobiology University Health Network University of Toronto #scGPT #GenerativeAI #AI4Science #Combio #opensource

Bo Wang

199,725 Aufrufe • vor 2 Jahren

How does an embryo reliably "compute" its form - "cell by cell" - using only local interactions and mechanics, yet produce a precise global body plan? I’m excited to share our Nature Methods paper "MultiCell: geometric learning in multicellular development", presenting #AIxBiology research led by Haiqian Yang and the result of a great collaboration with Ming Guo, George Roy, Tomer Stern, Anh Nguyen and Dapeng Bi. A long-standing challenge in developmental biology is to predict how thousands of cells collectively self-organize as tissues fold, divide, and rearrange. In MultiCell, we represent a developing embryo as a dual graph that unifies two complementary views of tissue mechanics with single-cell resolution: cells as moving points (granular) and cells as a connected foam (junction network). This lets the model learn dynamics from both geometry and cell–cell connectivity. On whole-embryo 4D light-sheet movies of Drosophila gastrulation (~5,000 cells), our model predicts key cell behaviors and the timing of events, including junction loss, rearrangements, and divisions with high accuracy, at single-cell resolution. Beyond prediction, the same representation supports robust time alignment across embryos and offers interpretable activation maps that highlight the morphogenetic "drivers" of development. The broader goal is a foundation for cell-by-cell forecasting in more complex tissues, and eventually for detecting subtle dynamical signatures of disease. Kudos to the team for this inspiring collaboration with brilliant researchers to push the boundary of AI for biology! Citation: Yang, H., Roy, G., Nguyen, A.Q., Buehler, M.J., et al. MultiCell: geometric learning in multicellular development. Nature Methods (2025), DOI: 10.1038/s41592-025-02983-x Code/data links are in the manuscript.

Markus J. Buehler

388,098 Aufrufe • vor 7 Monaten

Single-cell technologies now let us profile entire transcriptomes in individual cells. But how do we make sense of this complexity in a biologically meaningful way? Many methods summarise cells into a single embedding, but this often comes at the cost of interpretability, especially when multiple gene programs are active at once. We developed Tripso, a self-supervised transformer model that represents cells through multiple gene program-specific embeddings, while also uncovering new programs directly from the data. Instead of collapsing biology into a single vector, Tripso decomposes cell state into multiple representations, each reflecting a different gene program. We explored this across multiple systems. In human hematopoiesis, spanning development to aging, Tripso identified distinct age-associated program activity, including stronger JAK-STAT signalling in early life and dynamic IKZF1-related changes during B cell maturation. By comparing in vitro culture conditions with in vivo hematopoietic stem cell states, Tripso suggested that targeting the SEC61 translocon could enhance stem cell maintenance ex vivo, a prediction that we subsequently validated experimentally. In parallel, we identified a previously uncharacterised tissue-resident memory T-cell program associated with atopic dermatitis and mapped it to distinct spatial immune niches Together, these results show how modelling cells through gene programs can lead to interpretable and experimentally testable insights. More broadly, this work points toward a more interpretable and biologically grounded models of cell state. As single-cell datasets continue to grow, we hope approaches like Tripso will help bridge the gap between data-driven representations and biological insight. This work wouldn’t have been possible without the contributions of an amazing team. Thank you to co-first authors Marie, Tomoya Isobe, Amirhosein Vahidi, Carlo Leonardi, and everyone from roser's Lab, Haniffa Lab, Nicola Wilson and Bertie Gottgens's Lab, bringing together expertise across Cambridge Stem Cell Institute, Open Targets, Wellcome Sanger Institute and Cambridge University. Marie is one of the very best PhD students I have ever supervised. She is truly a force of nature, exceptionally resourceful, deeply innovative, and one of the most impressive scientists I have worked with. I am immensely proud of her and all that she has accomplished. As she begins her internship at Genentech , I have no doubt she will do amazing work there and continue to make her mark. paper: code:

Mo Lotfollahi

22,476 Aufrufe • vor 4 Monaten

Scientists just figured out how to reverse aging using AI. And this is a massive breakthrough. We can now reprogram any human cell back to age 20. Heart cells, brain cells, skin cells, all reset to their biological prime. And here’s the wildest part…the technology to do this, has already existed since 2012 (it won the Nobel Prize). But the real breakthrough wasn’t possible until this year, when they supercharged it with AI. It’s a wild story. So in 2006, scientists discovered Yamanaka factors. They’re proteins that can basically convert any normal cell into a universal stem cell. Now this was a huge deal, because these stem cells are basically like magic healers. If you have torn muscle tissue, you could inject these stem cells into the area and they will turn into the youthful muscle cells you need. So Yamanaka factors were this insane breakthrough, because they allowed any human to turn any cell you already have into these magic healers. But, there was one big problem… It turns out, the original Yamanaka factors weren’t very good at this stem cell conversion. They could do it, but they just weren’t very reliable. Enter OpenAI...and this is where things get crazy. OpenAI designed a special AI model built specifically to create new proteins. Think of it like ChatGPT but for protein engineering. So they took all the Yamanaka research and asked this new AI to go ham on improving it. And get this… Their version was 50x more effective than the original. They tested it on 50 year old cells and it successfully started repairing 30% of their cells in just 7 days. This is just science fiction…it actually happened. And it sounds crazy, but in a few years, humans will be able to take a shot that will literally reverse the age of their cells.

Whiplash347

68,643 Aufrufe • vor 9 Monaten

Google presents Still-Moving Customized Video Generation without Customized Video Data Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation is still in its infancy, primarily due to the lack of customized video data. In this work, we introduce Still-Moving, a novel generic framework for customizing a text-to-video (T2V) model, without requiring any customized video data. The framework applies to the prominent T2V design where the video model is built over a text-to-image (T2I) model (e.g., via inflation). We assume access to a customized version of the T2I model, trained only on still image data (e.g., using DreamBooth or StyleDrop). Naively plugging in the weights of the customized T2I model into the T2V model often leads to significant artifacts or insufficient adherence to the customization data. To overcome this issue, we train lightweight Spatial Adapters that adjust the features produced by the injected T2I layers. Importantly, our adapters are trained on "frozen videos" (i.e., repeated images), constructed from image samples generated by the customized T2I model. This training is facilitated by a novel Motion Adapter module, which allows us to train on such static videos while preserving the motion prior of the video model. At test time, we remove the Motion Adapter modules and leave in only the trained Spatial Adapters. This restores the motion prior of the T2V model while adhering to the spatial prior of the customized T2I model. We demonstrate the effectiveness of our approach on diverse tasks including personalized, stylized, and conditional generation. In all evaluated scenarios, our method seamlessly integrates the spatial prior of the customized T2I model with a motion prior supplied by the T2V model.

AK

40,485 Aufrufe • vor 2 Jahren