Video wird geladen...
Video konnte nicht geladen werden
Proteins can now talk. Introducing BioReason-Pro, the first reasoning model for protein function. A thread🧵
205,536 Aufrufe • vor 6 Monaten •via X (Twitter)
71 Kommentare

BioReason-Pro is a multimodal LLM that brings protein foundation models and LLMs together for reasoning.

It was trained on 130K+ protein reasoning traces and then refined further with RL.

It outperforms all prior methods in both Gene Ontology and free text prediction.

Even human experts preferred it over UniProt ground truth in 79% of the cases

BioReason-Pro correctly predicted a functional protein partner that was validated in a cryo-EM study. It's attention was right at the contact residues.

It has learned structural reasoning purely from trainnig. In a shocking case, when predicting protein scaffold activity, it attended to exactly the 3 residues out of thousands that had been repurposed from catalytic to scaffolding function.

You can talk to it here! Paper: Code: Data: Weights: , Catalogue of 240,000+ predictions:

This has been an incredible work of a big and powerful team! Thank you to Arman Seyed-Ahmadi (@arman1sa), Parsa Idehpour (@Radii2323), Omar Ibrahim, Purav Gupta, Jack Naimer, Kevin Zhu, Arnav Shah, Shihao Ma, Abhinav Adduri, Talu Güloglu, Nuo Liu, Haotian Cui, Arihant Jain, Max de Castro, Amirfaham Fallahpour, Antonio Cembellin-Prieto, John S. Stiles, Filip Nemčko, Alexander A. Nevue, Hyungseok C. Moon, Lucas Sosnick, Olivia Markham, Haonan Duan, Michelle Y. Y. Lee, Andrea F. M. Salvador, Chris J. Maddison, Christoph A. Thaiss, Chiara Ricci-Tam, Brian S. Plosky, Dave P. Burke (@davey_burke), Patrick D. Hsu (@pdhsu), Hani Goodarzi (@genophoria), and Bo Wang (@BoWang87). across Arc Institute (@arcinstitute), University Health Network (@UHN), Vector Institute (@VectorInst), University of Toronto (@UofT), Stanford University (@Stanford), and more!

Great project! I've been looking for someone to try this on protein function and your team did a great job!

Thank you Andrew, appreciate it!

the goat strikes again

🫡

summarize the paper

LMFAO

@adibvafa probably one of the most fascinating young researchers in ai x bio

🫡

Great work Adib, very interesting!

thank you Faraz!!

Thanks for dropping the code and the weights

of course!

Exceptional work. We're doing something similar...

super interesting!

Thank you!

very interesting work!

thank you!

Proteins talking is such a cool way to frame this! 🧬 What actually gets me excited here isn't just the technical achievement—it's what this means for people who need protein insights but can't access expensive labs or PhDs. When models like BioReason-Pro make complex biology more accessible, we're not just advancing science. We're democratizing it. Curious: do you think this will help rural clinics and smaller research teams catch up faster? Or will the tech still stay concentrated in big orgs?

no

lol fair! just genuinely excited about the accessibility angle though - when biology tools become more democratized, that's where real impact happens. not everyone needs a PhD to benefit from better protein insights!

yes!

Is that if i give him a sequence of RNA will predict which protein is closest to it. ?

It takes a protein sequence and target organism, it reasons what the protein does

I'm interested i will read the paper I want to make a post for it in my facebook page Thanks i will try it

awesome. congrats dude!!

Yoo thank you man!

Awesome! Congrats

Thanks Suraj!

this is a deal breaker for biology/life sciences students, researches and industrials! such an awesome idea

:D

this is dope!

Thank you!

Amazing work. Predicting those contact points is exactly what we are crowdsourcing right now for an "undruggable" cancer target. We have a $500k prize for whoever can computationally find a binder. Would love to see someone use BioReason-Pro to crack it!

Proteins talking now? Wild af

Great work! 110 pages. Is there metrics when InterPro domain is not found?

Yes its in supp figures (end of paper)

Yo brother

The logic is sound. But the real bottleneck isn't the model—it's the data quality.

Always!

woooow

woooow

110 page paper came as a surprise

@Ayush3241 lots of writing :))

The discussion on short peptide is really cool. RL model can effectively admit "I don't know" (even many trained human scientist cannot). Big congrats!!

Thank you! @Radii2323 truly cooked with RL

@Radii2323 Just out of curiosity: have you guys tested this on promiscuity? I don't if enough public data available out there to create a reasonable benchmark. But if it could work well, it can be quite useful for industrial applications.

@eigenron so sick!

@eigenron thank you :)

How does this compare to alphafold??

alphafold is for protein structure prediction, we are for protein function prediction

The key insight is bridging the representation gap — protein foundation models encode structural and evolutionary knowledge that LLMs alone can't capture, while LLMs bring reasoning and natural language interface. Training on 130K+ protein reasoning traces with RL refinement is a smart design choice, reminiscent of how reasoning models in other domains benefit from process reward signals. This could significantly lower the barrier for non-computational biologists in drug discovery to interrogate protein function directly.

Really interesting work. Curious whether BioReason-Pro surprised you more in benchmark performance or in the quality of its reasoning/explanations.

For me the most surprising result was its structural reasoning. We never taught it.

This is really cool, we should talk man!

For sure, DMs open!

Incredible 🫡

🫡

How could this help the world @grok?

wtf are we getting into

vibe proteining

amazing

🫡

@andrewwhite01 A Thread That Chase Can Grok (ATTCCG) @DeneckeChase

