Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Today, we're releasing the Carbon Annotation Database and the model behind it, Carbon-A. We've used this model to discover 566.34 million new gene candidates across 22,617 species, thousands of which have never been studied like this before. We wet-lab validated a number of these new genes in well-studied species,...

30,672 görüntüleme • 20 saat önce •via X (Twitter)

13 Yorum

Georgia Channing profil fotoğrafı
Georgia Channing20 saat önce

Carbon-A is an encoder with bi-directional attention that classifies nucleotides as coding (CDS) or non-coding in eukaryotes. We use a further small, post-processing model to identify genes from these nucleotide-level predictions. Carbon-A delivers a macro-F1 of 0.944, across mammals, other vertebrates, invertebrates, plants, fungi, and protozoa, when evaluated against RefSeq's reference genomes.

Georgia Channing profil fotoğrafı
Georgia Channing20 saat önce

We think these metrics may, in fact, be fully saturated, as the reference genome assemblies may have mistakes themselves. In discussions with the NCBI, we learned that the pig genome assembly that we trained with, released in 2017, was known to have many mistakes. Since we completed model training, a new and improved pig genome assembly has been released. Though Carbon-A had never seen the new assembly, it performed better on the newer assembly, likely having learned to overcome actual errors in the previous assembly.

Georgia Channing profil fotoğrafı
Georgia Channing20 saat önce

We show how our model identifies new genes through Iso-Seq experiments conducted, and we are excited to share results in the next few weeks on a genome never studied before.

Georgia Channing profil fotoğrafı
Georgia Channing20 saat önce

The Carbon Annotation Database makes available annotations for 22,617 species, which you can explore through our Carbon Annotation Database Explorer. This is a work in progress, and we will keep adding to the database until we have annotated all of GenBank—-so you should expect changes to our total counts over time.

Georgia Channing profil fotoğrafı
Georgia Channing20 saat önce

We hope that by releasing this to the community, we can create new levers for biologists to study the functional properties of many more species than possible before, and we look forward to feedback and improving this work over time. This is just one more step towards understanding life on earth, but there is still a huge amount of work to be done. We've identified millions of new genes, but we do not yet know what they do. We will need to work closely with the genomics community to turn these gene candidate discoveries into well-understood mechanisms useful for drug development, crop resilience, and more.

Georgia Channing profil fotoğrafı
Georgia Channing20 saat önce

Naturally, it's all an open-source collaboration. You can find the model, the training data, the annotations, and the technical report in the collection here: To get started with something lighter weight, read the blog:

Emma Scharfman profil fotoğrafı
Emma Scharfman20 saat önce

this is such a huge database that you're releasing! and it opens so many new doors, congrats 🧬🚀

Qiuyi Li profil fotoğrafı
Qiuyi Li20 saat önce

😍😍😍😍😍😍😍😍😍😍😍😍😍😍

英格兰AI哥 profil fotoğrafı
英格兰AI哥19 saat önce

@ClementDelangue 这些知识一定非常有意思,感谢AI加速人类多领域的演进

Julien Duquesne profil fotoğrafı
Julien Duquesne16 saat önce

Impressive! Congrats!!

Daniel Priscu profil fotoğrafı
Daniel Priscu19 saat önce

566 million candidates across 22k species is wild. how many of the wet lab checks came back clean?

Georgia Channing profil fotoğrafı
Georgia Channing18 saat önce

wdym?

Daniel Priscu profil fotoğrafı
Daniel Priscu17 saat önce

of the candidates you checked in the wet lab, what share held up as real genes?

Benzer Videolar

Welcome to the Lab of the Future! 🧬🤖 Excited to share LUMI-lab, out today in Cell — a self-driving platform that pairs an AI foundation model with a robotic lab to autonomously discover ionizable lipids (LNPs) for mRNA delivery. The core problem: Designing lipid nanoparticles (LNPs) is hard. The chemical space of ionizable lipids is vast, experimental cycles are slow, and — critically — historical LNP datasets are far too small to train a predictive model from scratch. Most AI approaches in this space hit a wall immediately: not enough data to learn from. Our solution: lab-in-the-loop foundation model learning. Instead of training on LNP data alone, LUMI starts as a transformer-based foundation model pretrained across broad chemical space, building rich molecular representations before it ever sees a single LNP experiment. Then it enters a closed loop with a robotic synthesis platform: predict → synthesize → assay → update. Each round of real wet-lab experiments fine-tunes the model, which then proposes smarter candidates for the next round. The lab isn't just validating AI predictions — it's actively teaching the model, continuously. What happened when we let it run: LUMI-lab autonomously synthesized and screened 1,700+ ionizable lipids in human bronchial epithelial cells. The top candidate — LUMI-6 — features a brominated lipid tail, a structural motif that had been largely overlooked in LNP design. LUMI found it without being told where to look. When formulated into LNPs and delivered intratracheally to mice, LUMI-6 achieved 20.3% gene editing efficiency in lung epithelial cells — a compelling result for one of the hardest-to-reach therapeutic targets, directly relevant to diseases like cystic fibrosis and alpha-1 antitrypsin deficiency. Why this matters beyond LNPs: This is a proof of concept for a broader thesis — that foundation model pretraining + active learning + robotic experimentation can overcome the data scarcity bottleneck that plagues AI-driven discovery in biology. You don't need a massive domain-specific dataset to start. You need a model that can generalize, a lab that can generate the right data, and a loop that connects them. Huge congratulations to first authors Yue Xu, Haotian Cui, and Kuan Pang, and to the entire Bowen LI team. Grateful to our collaborators at University Health Network and Leslie Dan Faculty of Pharmacy, and to Princess Margaret Cancer Centre Research Princess Margaret Cancer Centre Research. 📄 Paper:

Bo Wang

57,640 görüntüleme • 7 ay önce

Have you ever noticed what happens if a wild animal or plant disappears ? Most of us will never know. And even if we do, we rarely realise that our own survival and infact the survival of all life forms on Earth is deeply tied to species diversity. When a species disappears, it shakes this balance in ways we may not immediately see. Global assessments by IUCN Species Survival Commission WWF and UN Environment Programme warn that species are vanishing faster than ever. UNEP notes that nearly one million species are now at risk of extinction, many within decades. This is why species revival programmes matter. We need sustained efforts to rebuild populations of species facing decline before it is too late. In this context, Project Nilgiri Tahr assumes great significance as it aims to restore the population of this endangered iconic species endemic to Western Ghats. We just concluded the four-day population estimation exercise of the Nilgiri Tahr. It was inspiring to walk along with my hardworking TN Forest Department team across the rugged mountains and grasslands that are the Tahr’s abode, alongside Thiru Yash Veer IUCN India & Collector & DM, The Nilgiris. Using the Varudai app, we captured field data with greater precision under Project Nilgiri Tahr, now in its next phase with the Third synchronised survey across 14 divisions, 43 ranges, 124 beats, and 177 blocks, covering over 3,100 km with 800 staff. From 1,031 Tahrs in 2024 to 1,303 Tahrs in 2025, this iconic State Animal of Tamil Nadu is showing great signs of recovery.

Supriya Sahu IAS

11,663 görüntüleme • 5 ay önce