Loading video...

Video Failed to Load

Go Home

A protein language model predicts the effects of 450 million missense variants across >40,000 protein isoforms! Very cool work from Nadav Brandes Vasilis Ntranos Jimmie Ye in Nature Genetics

47,573 views • 3 years ago •via X (Twitter)

6 Comments

Jessica Andrade's profile picture
Jessica Andrade3 years ago

@BrandesNadav @vntranos @yimmieg @NatureGenet Very nice overview. I loved your content. You might also like:

Pat Adams's profile picture
Pat Adams3 years ago

@BrandesNadav @vntranos @yimmieg @NatureGenet You do an amazing job, conveying scientific information in an understandable way in these videos! Thank you!

ScienceStanley's profile picture
ScienceStanley3 years ago

@BrandesNadav @vntranos @yimmieg @NatureGenet Have worked quite a bit with protein language models at @RTTPatStanford mapping viral evolutionary paths... ...these models have a ways to go, but have enormous potential. Believe will be core part of the bioinformatician's tool box before too long 😇💜🙏

jamesjfink's profile picture
jamesjfink3 years ago

@BrandesNadav @vntranos @yimmieg @NatureGenet I love this!

Pat Adams's profile picture
Pat Adams3 years ago

@BrandesNadav @vntranos @yimmieg @NatureGenet Sequencing of Y chromosome. A story you could artfully communicate. Jackson Labs has a great article on this Thanks for all you do to contribute to making science accessible!

Nicolaas Van Renne's profile picture
Nicolaas Van Renne3 years ago

@BrandesNadav @vntranos @yimmieg @NatureGenet the only bad thing about these videos is that they should be two minutes with more info on the paper

Related Videos

What seemed like an intractable problem is now possible: To design proteins with a specified nonlinear mechanical response, capturing complex folding and unfolding mechanisms in singe and few-shot computations. We present ForceGen, an end-to-end algorithm for de novo protein generation based on nonlinear mechanical unfolding responses. Rooted in the physics of protein mechanics, this generative strategy provides a powerful way to design new proteins rapidly, including exquisite and rapid predictions about their dynamical behavior. Proteins, like any other mechanical object, respond to forces in peculiar ways. Think of the different response you'd get from pulling on a steel cable versus pulling on a rubber band, or the difference between honey and glass. Now, we can design proteins with a set of desirable mechanical characteristics, with applications from health to sustainable plastics. The key to solving this problem was to integrate a protein language model with denoising diffusion methods, and using accurate atomistic-level physical simulation data to endow the model a first-principles understanding. ForceGen can solve both forward and inverse tasks: In the forward task, we can predict how stable a protein is, how it will unfold and what the forces involved are, all given just the sequence of amino acids. In the inverse task, we can design new proteins that meet complex nonlinear mechanical signature targets. Read the paper, led by LAMM@MIT postdoc Bo Ni, published in Science Advances: Why do we care about the mechanics of proteins? The mechanics of proteins are critical elements of many living systems - as evidenced in many studies of mechanobiology. Through evolution, nature has presented a set of remarkable protein materials with unique mechanical functions like elastins, silks, keratins or collagens that play crucial roles in biology. However, going beyond natural designs to discover proteins that meet specified mechanical properties remains challenging. So far, the only way to do this was to use existing evolutionary concepts or to manually alter proteins. With our new generative model we can directly design proteins to meet complex nonlinear mechanical property-design objectives. ForceGen leverages deep knowledge on protein sequences from a pretrained protein language model and maps mechanical unfolding responses to create proteins. Via full-atom molecular simulations for direct validation from physical and chemical principles, we demonstrate that the designed proteins are de novo, and fulfill the targeted mechanical properties, including unfolding energy and mechanical strength, and a detailed unfolding force-separation curves. ForceGen offers rapid pathways to explore the enormous mechanobiological protein sequence space unconstrained by biological synthesis, to enable the discovery of new protein materials with superior mechanical properties. B. Ni, D.L. Kaplan, M.J. Buehler, ForceGen: End-to-end de novo protein generation based on nonlinear mechanical unfolding responses using a language diffusion model. Sci. Adv. 10, eadl4000 (2024). DOI: 10.1126/sciadv.adl4000 Codes and model weights available Hugging Face: David Kaplan

Markus J. Buehler

47,269 views • 2 years ago

Proof that Life is Intelligently Designed: Proteins. They do everything in your cells that keeps you alive. These microscopic molecules must have been created. Here is why: Proteins are the building blocks of all the little nano-machines that make up your cells. They do all the major work in your body - from creating energy to recycling waste. Proteins are made from amino acids connected into a specific sequence and then folded into a functional shape. You can think of a protein like a paragraph, and amino acids as the individual letters that spell out the words. Here is where it gets interesting... There are just 20 amino acids used in all of Life. Just 20 amino acids. Responsible for ~200 million unique proteins making up tens of thousands of cell types across Life. And scientists are constantly finding new ones. Scientists categorize proteins into Families. Protein Families are groups of proteins that share amino acid sequence similarities. There are ~22,000 Protein Families. Here's the part that screams they are designed... Protein Families have no evolutionary history. Even evolutionist scientists agree - they are absolutely unique. It's been said that proteins are like stars in a galaxy, and families are like galaxies, with vast empty space between them. Evolution is supposed to build by tweaking things that already exist. But no evolutionary history connects fundamentally distinct protein families. Experiments have been done to intentionally evolve one protein family into another. They fail every time. One protein family cannot be evolved into another. They cannot arise through evolution. But the math really makes it impossible... The odds of evolution finding even a single functional protein family is 1 chance in 10^77 possible amino acid sequences. That's a 1 with 77 zeroes. The odds of evolution finding 22,000 distinct protein families? Roughly 1 chance in 10^1,694,000. You don't need to be a mathematician to understand that this is impossible. Evolution works by random mutations tweaking what's already there. But protein families can't evolve from one another. And the math makes it absolutely absurd to believe. Monkeys banging on keyboards will never type out Shakespeare. And random mutations in sequences of amino acids could never create a single functional protein family. There is only one thing we know of that creates functionally specified sequences: Intelligence. Now here is the final cherry on top: Proteins are built by other proteins, assembled into a complex machine to do a job in cells. The assembly instructions for those proteins and the machines are found in DNA. But DNA requires protein machines for replication & repair. So DNA is required to make Proteins... ...but Proteins are required to make DNA *and* more Proteins. You can't have one without the other. They rely on each other. Which means one couldn't have evolved and then waited for the other. They couldn't do anything without each other. They had to be created at the same time to function together. Life was Divinely Designed. Proteins prove it.

Divinely Designed

18,886 views • 2 months ago

Prioritizing and interpreting disease-associated genetic variants remains one of the greatest challenges in human genetics. Today, we’re thrilled to introduce AlphaGenome Atlas 🧬, a genome-wide platform providing precomputed predictions for the regulatory effects of all ~9 billion possible single-letter changes and >100M observed indels in the human genome. Here is what Atlas delivers: 1. Variant Prioritization via AVI To prioritize variants, we developed the AlphaGenome Variant Impact (AVI) score. AVI predicts a unified score per variant, where higher values indicate greater disruption. It achieves state-of-the-art performance across diverse benchmarks. As a proof-of-concept with our collaborators at Broad Institute, AVI prioritized a deep-intronic variant in DNM1, helping solve a previously unexplained rare epileptic encephalopathy case by revealing a brain-specific cryptic splice site. 2. Multi-Layer Molecular Interpretation Variant prioritization is only half the battle; researchers also need to understand why a variant matters. Atlas decomposes variant effects across multiple interpretable layers: •Feature Attributions which decompose each variant’s score into specific biological modalities driving the impact. •Cell-Type Specificity: Precomputed predictions with AlphaGenome across hundreds of biosamples reveal the exact cellular context in which a variant acts. •Regulatory Grammar: Over 2,600 de novo DNA motifs (and >250B genome-wide instances) show when variants directly disrupt critical regulatory binding "words". Explore the resource: 🌐 Interactive browser & precomputed data: - 🎥 Video: - 📖 Blog: - 📄 Preprint:

Jun Cheng

11,596 views • 9 days ago

From a World-Renowned Oncologist: A Chilling Warning on Vaccines, Spike Protein, and the Children Now Dying from Once-Unheard-of Cancers Like Colon Cancer at Age 13. A chilling warning from Dr. Patrick Soon-Shiong: We are navigating a silent, post-pandemic health crisis, and the common culprit may be the spike protein. The renowned surgeon and scientist expresses profound alarm over the long-term effects he is now witnessing. The central question is no longer just the virus itself, but the impact of the spike protein—a component common to both the virus and the vaccine—which circulates in the bloodstream and permeates our tissues. This, he fears, is linked to a disturbing and unprecedented phenomenon: a super surge of cancers in young people. With a heavy heart, Dr. Soon-Shiong shares that in his decades in oncology, he has never witnessed what he sees now. He details the rise of aggressive, early-onset cancers, particularly colon cancer, in children as young as 10, 12, and 13 years old. He recounts the story of David Cohen of Oxford University, who was recently nearly in tears after a 13-year-old boy died of metastatic colon cancer in his care. "That's crazy," Dr. Soon-Shiong states, underscoring the abnormality of such a case. The crisis hit even closer to home in his own Los Angeles clinic. A child from Butler, Pennsylvania—the very town where President Trump was shot—came to him with metastatic pancreatic cancer after being turned away elsewhere. The reason? Hospitals declared it an "adult disease" they didn't know how to treat in a child. Despite efforts to get him care, by the time the family made the arduous journey from Butler to LA, the cancer had progressed too far. The child passed away. This is not just a statistic. It is a growing reality. Dr. Soon-Shiong's testimony is a urgent call to action, demanding we confront the complex legacy of the spike protein and investigate this alarming surge in young lives being cut short by diseases once thought exclusive to adulthood. The medical establishment must wake up. Our children's lives depend on it.

Camus

56,581 views • 11 months ago

A stunning testimony before the Massachusetts Legislature demands our full attention. Dr. Janci Chunn Lindsay, a PhD molecular biologist and toxicologist, delivered a methodical and devastating critique of the COVID-19 mRNA injections, asserting that the public was built upon a foundation of falsehoods. According to Dr. Lindsay, the very platform was "sold with a bunch of lies." She systematically listed the key promises made to the public that, in her professional assessment, have proven untrue: ▪️ The Nature of the Product: We were assured it was not a gene therapy. ▪️ Localized Effect: We were told the injection would remain localized in the arm. ▪ mRNA Stability: We were told the modified mRNA would break down rapidly. ▪️ Cellular Mechanism: We were told the spike protein could not enter the cell nucleus. ▪️ Genomic Impact: We were told reverse transcription and integration into human DNA was impossible. ▪️ Vaccine Purity: We were told there was no risk of DNA contamination altering our genetics. ▪️ Protein Specificity: We were told the shots would only produce the intended spike protein. ▪️ Treatment Alternatives: We were told no other safe and effective treatments existed. ▪️ Efficacy: We were told the shots were effective at preventing severe disease and death. ▪️ Transmission: We were told the products could not be shed to others. Her conclusion is stark: "All of these were lies." This testimony from a credentialed expert challenges the core narrative of the global pandemic response. It raises profound questions about institutional trust, scientific transparency, and informed consent. This is not a fringe theory; it is a formal allegation from a scientist with the background to make it. The conversation can no longer be avoided.

Camus

158,301 views • 11 months ago

Announcing How Transformer LLMs Work, created with Jay Alammar and Maarten Grootendorst, co-authors of the beautifully illustrated book, “Hands-On Large Language Models.” This course offers a deep dive into the inner workings of the transformer architecture that powers large language models (LLMs). The transformer architecture revolutionized generative AI; in fact, the "GPT" in ChatGPT stands for "Generative Pre-Trained Transformer." Originally introduced in the Google Brain team's groundbreaking 2017 paper "Attention Is All You Need," by Vaswani and others, transformers were a highly scalable model for machine translation tasks. Variants of this architecture now power today’s LLMs such as those from OpenAI, Google, Meta, Cohere, Anthropic and DeepSeek. In this course, you’ll learn in detail how LLMs process text. You'll also work through code examples that illustrate that transformer's individual components. In details, you’ll learn: - How the representation of language has evolved, from Bag-of-Words to Word2Vec embeddings to the transformer architecture that captures a word's meanings taking into account the context of other words in the input. - How inputs are broken down into tokens before they are sent to the language model. - The details of a transformer's main stages: Tokenization and embedding, the stack of transformer blocks, and the language model head. - The inner workings of the transformer block, including attention, which calculates relevance scores, and the feedforward layer, which incorporates stored information learned in training. - How cached calculations make transformers faster. - Some of the most recent ideas in the latest models such as Mixture-of-Experts (MoE) which uses multiple sub-models and a router on each layer to improve the quality of LLMs. By the end of this course, you’ll have a deep understanding of how LLMs actually process text and be able to read through papers describing the latest models and understand the details. Gaining this intuition will improve your approach to building LLM applications. Please sign up here:

Andrew Ng

259,920 views • 1 year ago

Demis Hassabis, the Nobel Prize winner who runs Google DeepMind just described the most consequential project on earth, and most people have no idea it exists. The project is called Isomorphic Labs and the goal is to end the way drugs have been developed for the last century. Here is the problem it is trying to solve. Developing a single drug today takes an average of 10 years, costs billions of dollars, and fails 90 percent of the time before it ever reaches a patient. Of every 10 drugs that enter clinical trials, only one makes it through. The other nine years of work, the other billions of dollars, the other scientific careers, gone. Hassabis believes AI can collapse that entire process from identifying a disease target to designing a compound that binds to it, predicts how it behaves in the body, and minimizes side effects , end to end, on a computer, before a single experiment is run. The foundation is AlphaFold, the AI system that solved one of biology's hardest problems predicting the 3D structure of every protein in the human body and won him the Nobel Prize in Chemistry in 2024. But knowing a protein's shape is only one part of designing a drug. Isomorphic is building what Hassabis describes as adjacent systems , AlphaFold 3, AlphaFold 4, and now a unified model called IsoDDE , that take the next steps. From designing the actual chemical compound that binds to the protein, predicting its binding strength, identifying new pockets to target that no one has ever found before. IsoDDE more than doubles the accuracy of AlphaFold 3 on the hardest protein-ligand prediction benchmarks that exist. Isomorphic is already running 18 to 19 live drug programs, cardiovascular disease, cancer, immunology in partnership with Eli Lilly, Novartis, and Johnson and Johnson. The first human clinical trial of a fully AI-designed drug is expected by the end of 2026. If that trial succeeds, it will be the first time in history that a drug put into a human body was designed not by a team of chemists working for a decade but by an AI working for months. Hassabis's long-term vision is even more direct, one day you describe a disease, click a button, and a drug blueprint comes out the other side. AI will solve almost all diseases within 10 years.

Milk Road AI

36,062 views • 5 months ago