Загрузка видео...

Не удалось загрузить видео

На главную

🚨 Just released: MIRIAD, a million-scale medical QA dataset to ground LLMs in reliable medical knowledge. 5.8M question-answer pairs, each distilled from peer-reviewed literature! 🔥 That's structured, high-quality data built for medical AI. 🧵 ↓

48,127 просмотров • 1 год назад •via X (Twitter)

Комментарии: 12

Фото профиля Charly Wargnier
Charly Wargnier1 год назад

Why this matters: Most RAGs or agentic systems rely on raw text: messy, noisy, and poorly aligned with downstream tasks. MIRIAD offers an alternative: ↳ Curated QA pairs ↳ Structured for retrieval ↳ Optimized for grounding and evaluation

Фото профиля Charly Wargnier
Charly Wargnier1 год назад

What you can do with MIRIAD: - Boost RAG/agentic system performance - Train medical retrievers - Detect hallucinations in clinical LLMs - Build knowledge-grounded expert assistants

Фото профиля Charly Wargnier
Charly Wargnier1 год назад

Also included: MIRIAD Atlas, an interactive map of 56 specialties that lets you explore, search, and trace answers back to the original source literature. Semantic search built in! 🔥

Фото профиля Charly Wargnier
Charly Wargnier1 год назад

The dataset was built through a semi-automated pipeline: LLM rephrasing → automatic filtering → multi-level QC → expert human annotation

Фото профиля Charly Wargnier
Charly Wargnier1 год назад

🤗 Dataset on @huggingface:

Фото профиля Charly Wargnier
Charly Wargnier1 год назад

🔗 Website:

Фото профиля Charly Wargnier
Charly Wargnier1 год назад

📜 Preprint:

Фото профиля Charly Wargnier
Charly Wargnier1 год назад

💻 Code:

Фото профиля Charly Wargnier
Charly Wargnier1 год назад

🪐 Demo:

Фото профиля Charly Wargnier
Charly Wargnier1 год назад

Credits for this great work! Supervisor: @Michael_D_Moor Co authors: Qinyue Zheng @qinzytech Salman Abdullah @salmanabdullah_ Sam Rawal @samarthrawal Cyril Zakka, MD @cyrilzakka Sophie Ostmeier @SophieOstmeier Eric Topol @EricTopol Maximilian Purk Eduardo P. Reis Jure Leskovec

Фото профиля Charly Wargnier
Charly Wargnier1 год назад

If you found this helpful, a like or RT goes a long way for this open source project to be discovered and used by many! Follow me → @datachaz for insights on lesser known projects like this and more content on LLMs, AI agents, and data science 🦾

Фото профиля Naveen Sankar S
Naveen Sankar S1 год назад

🌱 The gut-lung axis may hold the key to managing asthma, COPD, and ARDS. Explore how gut health may influence respiratory diseases and the future of targeted therapies. Check out the latest research at #GutLungAxis #Microbiome #LungHealth #health #news

Похожие видео

🎉 The best way to start the week is to find out that our MedSAM is finally published today in Nature Communications! **Segment anything in medical images** Paper: arXiv: Data & Code: MedSAM is the first promotable foundation model for medical image segmentation. **Highlights**: ⭐ Before its formal publication, we have received 220 citations and 1400+ GitHub stars 🙏🙏❤️‍🔥❤️‍🔥❤️‍🔥 📊 We curated a large-scale medical image dataset with 1,570,263 image-mask pairs, covering 10 imaging modalities and over 30 cancer types. 🚀 Built on top of SAM (AI at Meta ) with transfer learning, we have significantly enhanced its segmentation performance of medical images. 📈 Comprehensive evaluations of 86 internal validation tasks and 60 external validation tasks demonstrate its better accuracy and robustness than modality-wise specialist models. **What is Next? --- Clinical Translation!!** 🍕Our next goal is to make the model deployable on laptops (CPUs) or other edge devices without reliance on GPUs. We have distilled a lightweight model, LiteMedSAM, offering a speed boost of 10x while maintaining accuracy. Plus, we have integrated it into the 3D Slicer plugin, providing an efficient tool for medical image segmentation. 🌐 To further promote developments in this field, we organize a competition on #CVPR2026: Segment Anything in Medical Images on Laptop! An out-of-the-box baseline has been released to reduce the entry barriers. Welcome to join us to push the boundary further: 🙏 Massive thanks to MetaAI AI at Meta for their open-source project SAM and many reviewers/users for their invaluable feedback. A huge shoutout to my postdoc Jun Ma (JunMa) for his leadership on this project!! UHN AI Hub Vector Institute Peter Munk Cardiac Centre AI Department of Laboratory Medicine & Pathobiology U of T Department of Computer Science University of Toronto University Health Network Brad Wouters 🇨🇦 Barry Rubin MD, PhD, FRCSC Shaf Keshavjee

Bo Wang

140,208 просмотров • 2 лет назад

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 просмотров • 3 месяцев назад

Scale alone is not enough for AI data. Quality and complexity are equally critical. Excited to support all of these for LLM developers with Snorkel AI Data-as-a-Service, and to share our new leaderboard! — Our decade-plus of research and work in AI data has a simple point: scale alone is not enough. AI success is all about the quality, complexity, and distribution of data—in addition to volume. We’re excited to be powering leading LLM developers with Snorkel AI Expert Data-as-a-Service, our white glove service for custom, expert-level AI datasets—and to now preview some of what we’re building via our new Expert Data Leaderboard (🔗 in 🧵) + upcoming OSS dataset releases! Snorkel Expert Data-as-a-Service is built to meet the rapidly evolving data needs of the agentic AI world—where success is built on the quality, complexity, and distribution of datasets, in addition to size and scale. This kind of high-quality, frontier AI data can only come from a union of technology and human expertise. With Snorkel Expert Data-as-a-Service, we’re powering frontier LLM developers across agentic, expert knowledge, reasoning, coding, multi-modal, and other task types via the combination of these two key components: - (1) The Snorkel Expert Network: A global team of subject matter experts focused wholly on specialized knowledge–spanning thousands of topics in STEM/academic, vertical/professional, and consumer/lifestyle domains. - (2) Snorkel AI Data Development Platform: Our unique programmatic data curation and quality control platform, accelerating and improving expert authoring and review through principled techniques developed over the last decade of R&D. Now: we’re incredibly excited to showcase some of the power of Snorkel Expert Data-as-a-Service via the new Snorkel Leaderboard—putting frontier models to the test in complex, agentic, and reasoning settings inspired by real industry scenarios (not esoteric puzzles)! We’ll be releasing new leaderboards and accompanying expert-verified open source datasets (coming soon!) regularly. To start, we’re sharing three initial ones in preview: - SnorkelFinance: Q&A over financial documents requiring agentic tool-calling and reasoning - SnorkelUnderwrite: Agentic insurance tasks requiring industry-specific reasoning and tool use - SnorkelSequences: Mathematical tasks requiring compositional multi-step reasoning

Alex Ratner

495,851 просмотров • 1 год назад

Studies have shown ChatGPT outperforms human annotators for Structured Data by about 25% and costs 30x less. 1 In just 2 months, miners on SN33 running ChatGPT without optimization can’t survive. Today we announce SN33 is now ReadyAI to fully align with our mission 👇 SN33 is building a more performant and significantly cheaper alternative to Scale AI Today structured data is performed primarily by human annotation services like Amazon’s Mechanical Turk and Scale AI It is now more important than ever for every business and individual to make their data AI Ready. However, taking unstructured data and making it Structured Data using today’s tools is extremely costly. SN33 revolutionizes this process, unlocking immense opportunities for commercialization. We lay out the vision for it in this detailed blog post: Validators TODAY can monetize access to this structured data pipeline independently, but we’re streamlining this process, launching a frontend soon that any validator can opt into to provide bandwidth. We've received great feedback from the community, recognizing that what we're building goes far beyond Conversational AI. Building the world's largest annotated conversational dataset (which we've already accomplished) is just one of countless real-world applications for SN33's Structured Data pipeline. We're building a decentralized Scale AI, offering a full suite of Structured Data commodities—from text metadata tagging (available today) to fully customizable queries for company-specific data annotation use cases and image metadata tagging coming soon 👀. Thanks for all the feedback! It has been invaluable so keep bringing it to us! 🙏$TAO Openτensor Foundaτion 1 “ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks” shows “The zero-shot accuracy of ChatGPT exceeds that of crowd-workers by about 25 percentage points on average [...] Moreover, the per-annotation cost of ChatGPT is less than $0.003—about thirty times cheaper than MTurk”

David Fields

13,639 просмотров • 1 год назад