Загрузка видео...
Не удалось загрузить видео
Today, we're releasing the Carbon Annotation Database and the model behind it, Carbon-A. We've used this model to discover 566.34 million new gene candidates across 22,617 species, thousands of which have never been studied like this before. We wet-lab validated a number of these new genes in well-studied species,... show more
30,672 просмотров • 20 часов назад •via X (Twitter)
Комментарии: 13

Carbon-A is an encoder with bi-directional attention that classifies nucleotides as coding (CDS) or non-coding in eukaryotes. We use a further small, post-processing model to identify genes from these nucleotide-level predictions. Carbon-A delivers a macro-F1 of 0.944, across mammals, other vertebrates, invertebrates, plants, fungi, and protozoa, when evaluated against RefSeq's reference genomes.

We think these metrics may, in fact, be fully saturated, as the reference genome assemblies may have mistakes themselves. In discussions with the NCBI, we learned that the pig genome assembly that we trained with, released in 2017, was known to have many mistakes. Since we completed model training, a new and improved pig genome assembly has been released. Though Carbon-A had never seen the new assembly, it performed better on the newer assembly, likely having learned to overcome actual errors in the previous assembly.

We show how our model identifies new genes through Iso-Seq experiments conducted, and we are excited to share results in the next few weeks on a genome never studied before.

The Carbon Annotation Database makes available annotations for 22,617 species, which you can explore through our Carbon Annotation Database Explorer. This is a work in progress, and we will keep adding to the database until we have annotated all of GenBank—-so you should expect changes to our total counts over time.

We hope that by releasing this to the community, we can create new levers for biologists to study the functional properties of many more species than possible before, and we look forward to feedback and improving this work over time. This is just one more step towards understanding life on earth, but there is still a huge amount of work to be done. We've identified millions of new genes, but we do not yet know what they do. We will need to work closely with the genomics community to turn these gene candidate discoveries into well-understood mechanisms useful for drug development, crop resilience, and more.

Naturally, it's all an open-source collaboration. You can find the model, the training data, the annotations, and the technical report in the collection here: To get started with something lighter weight, read the blog:

this is such a huge database that you're releasing! and it opens so many new doors, congrats 🧬🚀

😍😍😍😍😍😍😍😍😍😍😍😍😍😍

@ClementDelangue 这些知识一定非常有意思,感谢AI加速人类多领域的演进

Impressive! Congrats!!

566 million candidates across 22k species is wild. how many of the wet lab checks came back clean?

wdym?

of the candidates you checked in the wet lab, what share held up as real genes?
