Video yükleniyor...
Video Yüklenemedi
🤗🤗🤗introducing Hugging Science -- the home of AI for science 🤗🤗🤗 open models and datasets are the powerhouse of science (see the PDB), but finding the models and data you actually need for your breakthrough is hard af you shouldn't need to scrape arxiv, own your own wetlab, fight... show more
226,376 görüntüleme • 5 ay önce •via X (Twitter)
50 Yorum

This is exactly the move. Discoverability of real scientific data + models on the Hub has been the missing piece for years. OpenMed's 1,000+ medical models all sit on top of @huggingface infra. becoming its own destination is overdue and welcome. 👏

@lvwerra Nice, more incoming from our side

@lvwerra OOOooooOOOo say more?

@lvwerra Ahaha more in counting!!

Great resource! Hope it expands beyond biology to physics (i.e. Materials discovery).

Yeah, tons and tons of materials data and models -- not to mention multiple benchmarks!

Yaaaaaaay 🚀🚀🚀🚀

Excellent initiative thank you @cgeorgiaw !

😊😊😊

Dang, that's incredible! I see potential for expansion. Do you envision this to eventually grow into a science data repository for *everything*?

We want it to be comprehensive but not overwhelming. A huge issue with scientific data is that you have tons and tons of random stuff where you're not sure what it is and where it came from. We are trying to mitigate that issue by putting forward really high quality, well-structured datasets that you can build with this instant.

Makes sense. Great work, that's going to help a lot of researchers.

What do you think is the best large language model for molecular biology questions and projects?

nice! I don't see and CFD datasets in the physics or engineering sections?

Been thinking about including stuff like this: But would love more suggestions. In general trying to include things that are sufficiently well-documented that people could build without a huge deep dive.

There all have dataset cards to some extent or another have a few more in there too

Epic will add TYSM!!!

Congrats on the launch. Can I pitch a challenge and what does standing one up look like? The one I'd love to see: pre-symptomatic Parkinson's detection from voice. Biomarkers predict PD 3-5 years early. Consented speech corpus + biomarker model + years-to-symptom leaderboard.

Yeah, 100%! We have a blog on the site about how to set up a challenge, but it’s pretty straightforward. You basically need a frontend leaderboard and an eval metric(s) that you can calculate from what people submit. You can also clone the set-ups of previous challenges that we’ve done. Once you’ve got it, reach out!

@prasithg I have the same question. Can you point me to the blogpost please?

@prasithg This blog is very how-to oriented, but if you feed your agent of choice with this blog, your idea, and an eval metric it should be able to set up a challenge without an issue.

I have such immense respect for you guys this is HUGE

Hey @atranscendedman check this out

Congratulations!

@ayirpelle Exciting times ahead 😁

This is just great, Hugging Science, I just love it. I have many suggestions eg bacterial datasets downloaded paper by paper that have been accumulated on various hard drives. Have already spotted databases I want to use. Thank you!

I have been trying to upload models and datasets for a while now ....now huge upgrade

Science deserves its own Spotify Wrapped moment for datasets.

😂😂😂

Cool stuff appears on X everyday

🤗x🧬 = ✅

What's the typical workflow you're envisioning? Like researcher searches "protein folding transformers" and gets ranked models with actual performance metrics instead of digging through 50 papers to find which ones have usable code?

That would be sick -- need a leaderboard for that. Even simpler workflow is "I'm trying to train a DNA model, what is the set of high-quality, known-provenance datasets I can use off the bat?" so you can go from idea to MVP way, way faster

科学のための人工知能のホーム、Hugging Scienceを紹介します。開かれたモデルとデータセットは科学の原動力です。しかしながら、実際にあなたのブレークスルーに必要なモデルやデータを見つけるのは本当に難しいですよね。しかし、私たちがそれを変えています。すべての最高の科学を一か所に集めました。さまざまなデータが準備されており、トレーニングが可能です。さらに、ドメイン、タスク、キーワードでフィルタリングや検索もできます。科学の進展に貢献しましょう!

@_lewtun > 11TB of PDEs Well there goes sleep and weekends.

This is awesome! I'm gonna save this for superlab :D

@ClementDelangue did you approve this?

i need DFT and NBO analysis on molecules in aqueous solution!

Woah

honestly the discovery problem is real, my main question is how did you determine "best" vs just having everything - because hf already has a lot of this stuff and the filtering is the actual pain point. like is this curated or just aggregated

curated! we selected for extremely well-documented, extensive, known-provenance datasets and models to minimize the "what even is this???" feeling. if you have any more suggestions how we could do that better, would love that (or just like specific pains too)

stop building stellarators when you can just scrape the data and bill it to your investors honest truth, most labs are just

Congrats!

"Open models and datasets are crucial, but sometimes scraping ArXiv or other sources is necessary for research. It's not about owning everything, but rather having access to what's out there, even if it means some extra legwork."

@Quetzally_Med

Hugging face is so convenient

Long live humanity!

Thank you for this gift to the community 🙏

AWESOME NEW AGE KAGGLE

@_akhaliq What can I build with this? Tell me!!

