Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

1/10 image-based screening advanced drug discovery but scaling to massive perturbation space is hard! Given a cell image, we asked if we could predict the morphological effect induced by a perturbation! Led by Alessandro Palma & w Fabian Theis we propose IMPA

40,670 Aufrufe • vor 3 Jahren •via X (Twitter)

10 Kommentare

Profilbild von Mo Lotfollahi
Mo Lotfollahivor 3 Jahren

2/n: IMPA uses "style-transfer" to model phenotypic responses in high-content imaging. It decomposes images into perturbation styles and cell contents, generating counterfactual images of perturbed cells.

Profilbild von Mo Lotfollahi
Mo Lotfollahivor 3 Jahren

3/n: Our model employs an autoencoder with versatile perturbation embeddings derived from pre-trained molecules or gene embedding. IMPA trains on adversarial methods to learn perturbation-specific styles for precise image generation.

Profilbild von Mo Lotfollahi
Mo Lotfollahivor 3 Jahren

4/n: IMPA on BBBC021 dataset images of perturbed MCF7 cells. Using RDKit drug embeddings, our model predicts morphological changes while preserving content. IMPA accurately generates actin degradation and nuclear integrity loss with Vincristine and Cytochalasin B.

Profilbild von Mo Lotfollahi
Mo Lotfollahivor 3 Jahren

5/n: CellProfiler (@DrAnneCarpenter) reveals IMPA's ability to capture relevant morphological features, outperforming control cells. The model generates perturbation-specific images that closely match morphological shifts in actual perturbed images.

Profilbild von Mo Lotfollahi
Mo Lotfollahivor 3 Jahren

6/n: IMPA compares well with existing GAN models for style transfer, exhibiting promising performance in evaluation metrics such as FID, density, and coverage, while maintaining a competitive mode of action generation accuracy.

Profilbild von Mo Lotfollahi
Mo Lotfollahivor 3 Jahren

7/n: IMPA's smooth style space allows for drug response interpolation. This allows navigation of the style space and studying the associated morphological response and effect similarities between perturbations.

Profilbild von Mo Lotfollahi
Mo Lotfollahivor 3 Jahren

8/n: IMPA predicts responses to unseen drugs when they are structurally similar to training compounds with the same mode of action. However, IMPA’s accuracy drops when unseen treatments are not structurally related/functionally similar training drugs.

Profilbild von Mo Lotfollahi
Mo Lotfollahivor 3 Jahren

9/n: We compared IMPA to Mol2Image, showcasing distinct advantages. While Mol2Image generates images from pure noise, IMPA performs style transfer on existing images, enhancing the study of differential morphology and significantly improving overall performance.

Profilbild von Mo Lotfollahi
Mo Lotfollahivor 3 Jahren

10/n: We applied IMPA to predict gene KO effects on two datasets (BBBC025, RxRx1), using Gene2Vec for embeddings. We showed that the model captures morphological changes caused by active perturbations, correctly identifying subtle changes in less active phenotypes!

Profilbild von Fabian Theis
Fabian Theisvor 3 Jahren

@ale__palmaa Happy that our „image perturbation autoencoder“ approach is out: Led by @ale__palmaa and @mo_lotfollahi, we learn the effect of drug/CRISPR perturbations on cell morphometry. This lets you style-transfer cell images to those under various perturbations.

Ähnliche Videos

Excited to share our new work. Over the past decade, single-cell genomics has transformed our ability to map cellular systems. But a major question remains: Can we predict how perturbations reshape cellular trajectories over time? In 2018, we first showed that it is possible to predict cellular responses to perturbations — ranging from disease signals to chemical treatments — even in unseen contexts. In 2022, we introduced CPA (MSB 2022; NeurIPS 2022), extending this idea to predict responses to unseen chemical and genetic perturbations, including their combinations. Since then, the field of perturbation modeling has grown enormously. The community has pushed the space forward with many creative ideas and powerful models. It’s exciting to see how fast things are moving — even though many fundamental challenges remain. One of the biggest is that cells are not static. They move through trajectories during development, immune responses, and disease. Yet most current models still predict perturbation effects within a single state, rather than how early perturbations propagate across future states and reshape downstream outcomes. To address this, we developed PerturbGen, a trajectory-aware generative AI model that predicts how genetic perturbations reshape downstream cellular states. Huge credit to the people who made this work possible. Thanks to co-first authors Kevin Ly, Adib Miraki, Tomoya Isobe, AmirHoss3in Vahidi, Delshad Vaghari & Anthony Rostron. Special recognition to Kevin Ly and Adib Miraki for driving this work over the finish line. Grateful for our outstanding collaborators from Haniffa Lab, Bertie Gottgens lab Gosia Trynka and many others — a true cross-institute effort across Cambridge Stem Cell Institute, Open Targets ,Wellcome Sanger Institute and Cambridge University.🎉 PerturbGen learns transcriptional dynamics across cellular trajectories. By introducing perturbations at an early source state, it can simulate how these effects propagate into future states along differentiation trajectories. Scaling this across genes enables the creation of dynamic in silico perturbation atlases — maps of how perturbations reshape biological trajectories over time. We explored this idea across three biological questions. First, in a human in vivo LPS immune challenge, PerturbGen predicted that perturbing a transient IL1B signal dampens downstream inflammatory programs in myeloid cells, with pathway changes reversing signatures observed in an independent IL-1β stimulation experiment. Second, in human hematopoiesis, PerturbGen predicted transcriptional responses to CRISPR transcription factor knockouts and enabled construction of perturbation atlases revealing lineage- and age-specific regulatory programs. These programs could also be linked to human genetics and blood diseases, including recapitulation of signatures associated with ETV6-related thrombocytopenia. Finally, we asked whether perturbation modeling could help improve complex tissue models. We built a dynamic perturbation atlas of human skin organoids to identify perturbations that could guideorganoid cells towardhuman fetal skin states. PerturbGen prioritized activation of Wnt signaling via GSK3β inhibition. Experimental validation confirmed the prediction: treatment with CHIR99021 induced stromal gene programs and shifted organoid fibroblasts toward transcriptional states observed in fetal skin stroma. Together, these results show how trajectory-aware perturbation modeling can connect gene perturbations to developmental programs, human genetics, disease mechanisms, and experimental interventions. More broadly, we think these point toward a future where single-cell atlases become predictive systems. As atlases expand across tissues, developmental windows, and modalities, models like PerturbGen could enable dynamic, virtual perturbation atlases— allowing us to simulate interventions, generate hypotheses, and design experiments before stepping into the lab. Preprint Code Excited to see how the community builds on this work.

Mo Lotfollahi

17,157 Aufrufe • vor 5 Monaten

[SIGGRAPH 2025] Photoreal Scene Reconstruction from an Egocentric Device Contributions: 1. We address the importance of employing visual-inertial bundle adjustment (VIBA) that accounts for the rolling-shutter behavior of the RGB camera. This provides a continuous camera trajectory to model pixel movement in neural reconstruction. Our experiments demonstrate that using VIBA consistently improves the novel view quality in Gaussian Splatting by +1 dB in PSNR. 2. We introduce a rasterization-based image formulation pipeline that addresses common artifacts in physical image formation, including rolling shutter, lens shading, exposure, and gain compensation. Our approach is distinct in that we represent image poses as posed pixel arrays sampled from a continuous trajectory, rather than assigning a single camera pose per image, and preserve the merit of Gaussian rasterization. Unlike existing methods that require ray-tracing Gaussians, e.g., [Moenne-Loccoz et al. 2024], our formulation is applicable to general-purpose rasterization-based Gaussian splatting. When applied to 3D Gaussian Splatting (3DGS) [Kerbl et al. 2023], our approach can further enhance reconstruction quality by +1 dB. We outperform existing baselines and demonstrate a substantial quality improvement in handling complex scenes observed by egocentric devices. 3. To reduce the effect of blur from rapid head motion in darker indoor scenes, we propose a strategy of deliberately underexposing input videos during capture, inspired by HDR+ [Hasinoff et al. 2016]. We demonstrate that we can reconstruct high-quality, noise-free scene radiance from noisy, dim input videos, and further render sharp, blur-free videos at a higher dynamic range.

MrNeRF

15,244 Aufrufe • vor 1 Jahr

[CLIP] by Hand ✍️ The CLIP (Contrastive Language–Image Pre-training) model, a groundbreaking work by OpenAI, redefines the intersection of computer vision and natural language processing. It is the basis of all the multi-modal foundation models we see today. How does CLIP work? Goal: 🟨 Learn a shared embedding space for text and image [1] Given ↳ A mini batch of 3 text-image pairs ↳ OpenAI used 400 million text-image pairs to train its original CLIP model. Process 1st pair: "big table" [2] 🟪 Text → 2 Vectors (3D) ↳ Look up word embedding vectors using word2vec. [3] 🟩 Image → 2 Vectors (4D) ↳ Divide the image into two patches. ↳ Flatten each patch [4] Process other pairs ↳ Repeat [2]-[3] [5] 🟪 Text Encoder & 🟩 Image Encoder ↳ Encode input vectors into feature vectors ↳ Here, both encoders are simple one layer perceptron (linear + ReLU) ↳ In practice, the encoders are usually transformer models. [6] 🟪 🟩 Mean Pooling: 2 → 1 vector ↳ Average 2 feature vectors into a single vector by averaging across the columns ↳ The goal is to have one vector to represent each image or text [7] 🟪 🟩 -> 🟨 Projection ↳ Note that the text and image feature vectors from the encoders have different dimensions (3D vs. 4D). ↳ Use a linear layer to project image and text vectors to a 2D shared embedding space. 🏋️ Contrastive Pre-training 🏋️ [8] Prepare for MatMul ↳ Copy text vectors (T1,T2,T3) ↳ Copy the transpose of image vectors (I1,I2,I3) ↳ They are all in the 2D shared embedding space. [9] 🟦 MatMul ↳ Multiply T and I matrices. ↳ This is equivalent to taking dot product between every pair of image and text vectors. ↳ The purpose is to use dot product to estimate the similarity between a pair of image-text. [10] 🟦 Softmax: e^x ↳ Raise e to the power of the number in each cell ↳ To simplify hand calculation, we approximate e^□ with 3^□. [11] 🟦 Softmax: ∑ ↳ Sum each row for 🟩 image→🟪 text ↳ Sum each column for 🟪 text→ 🟩 image [12] 🟦 Softmax: 1 / sum ↳ Divide each element by the column sum to obtain a similarity matrix for 🟪 text→🟩 image ↳ Divide each element by the row sum to obtain a similarity matrix for 🟩 image→🟪 text [13] 🟥 Loss Gradients ↳ The "Targets" for the similarity matrices are Identity Matrices. ↳ Why? If I and T come from the same pair (i=j), we want the highest value, which is 1, and 0 otherwise. ↳ Apply the simple equation of [Similarity - Target] to compute gradients of for both directions. ↳ Why so simple? Because when Softmax and Cross-Entropy Loss are used together, the math magically works out that way. ↳ These gradients kick off the backpropagation process to update weights and biases of the encoders and projection layers (red borders).

Tom Yeh

67,874 Aufrufe • vor 2 Jahren

CLIP by hand ✍️ ~ 13 steps walkthrough below CLIP, Contrastive Language-Image Pre-training, is OpenAI's answer to a question that sounds impossible: how do you put a sentence and a picture in the same space? CLIP shipped when OpenAI was still open, and those embeddings were shared far and wide. Almost every multimodal model you use today descends from them. How does it work? Goal: learn one shared embedding space for text and images. = 1. Given = A mini batch of three text-image pairs. OpenAI trained the original on 400 million. = 2. Text to vectors = Let us look up each word with word2vec. = 3. Image to vectors = We cut each image into two patches and flatten them. Now text and pixels are both just numbers. = 4. The other pairs = Repeat steps 2 and 3 for the rest of the batch. = 5. Encode = Let us push both sides through their encoders, a linear layer and a ReLU. In practice these are transformers, but the shape of the operation is the same. = 6. Mean pooling = We average across the columns, so each image and each sentence collapses to a single vector. = 7. Projection = The text vectors are 3D and the image vectors are 4D, so they cannot be compared at all. A linear layer projects both to 2D. That 2D space is the shared embedding space, and getting here is the whole point of the model. = 8. Prepare for matmul = Let us copy the text vectors down and the transposed image vectors across. = 9. MatMul = We multiply, which takes the dot product of every text vector with every image vector. Each cell is one estimate of how well a sentence matches a picture. = 10. Softmax, e to the power = Raise e to each cell. To keep it hand sized we approximate e with 3. = 11. Softmax, sum = Sum each row for image to text, each column for text to image. = 12. Softmax, normalize = Divide, and out come two similarity matrices, one per direction. = 13. Loss gradients = The targets are identity matrices: a pair that belongs together should score 1, every other cell 0. Subtract the target from the similarity and you have the gradients, in both directions. The takeaway: pairing a picture with a sentence comes down to a single dot product. Everything before step 9 is the work of getting them into one shared space, so that the dot product finally means something. 💾 Save this post!

Tom Yeh

20,750 Aufrufe • vor 1 Monat

VideoRF: Rendering Dynamic Radiance Fields as 2D Feature Video Streams paper page: Neural Radiance Fields (NeRFs) excel in photorealistically rendering static scenes. However, rendering dynamic, long-duration radiance fields on ubiquitous devices remains challenging, due to data storage and computational constraints. In this paper, we introduce VideoRF, the first approach to enable real-time streaming and rendering of dynamic radiance fields on mobile platforms. At the core is a serialized 2D feature image stream representing the 4D radiance field all in one. We introduce a tailored training scheme directly applied to this 2D domain to impose the temporal and spatial redundancy of the feature image stream. By leveraging the redundancy, we show that the feature image stream can be efficiently compressed by 2D video codecs, which allows us to exploit video hardware accelerators to achieve real-time decoding. On the other hand, based on the feature image stream, we propose a novel rendering pipeline for VideoRF, which has specialized space mappings to query radiance properties efficiently. Paired with a deferred shading model, VideoRF has the capability of real-time rendering on mobile devices thanks to its efficiency. We have developed a real-time interactive player that enables online streaming and rendering of dynamic scenes, offering a seamless and immersive free-viewpoint experience across a range of devices, from desktops to mobile phones.

AK

38,686 Aufrufe • vor 2 Jahren