Loading video...

Video Failed to Load

Go Home

AI has become a useful research partner. It can find information, summarize evidence, and suggest ideas. But can it help us think better? Can it help us navigate the complex nuances of high-stakes decisions? Today, we’re taking the next step toward our mission of scaling up good reasoning for...

10,307 views • 1 month ago •via X (Twitter)

7 Comments

Elicit's profile picture
Elicit1 month ago

We built the Research Agent around five capabilities: Capability 1: A harness with reasoning principles Reason with greater confidence with built-in models trained on BioDecisionBench, a benchmark designed to evaluate reasoning quality and handle uncertainty around decisions in pharma.

Elicit's profile picture
Elicit1 month ago

Capability 2: Diverse data sources Make decisions based on the full picture by synthesizing evidence across scientific literature, clinical trials, patents, regulatory and commercial sources, the web, and your own internal data.

Elicit's profile picture
Elicit1 month ago

Capability 3: Decision-ready outputs Move from research to decisions in a single workflow by turning evidence into fully cited reports, tables, analyses, visualizations, slides, and documents.

Elicit's profile picture
Elicit1 month ago

Capability 4: Transparency by default Trust and defend every decision with transparent methods trail and sentence-level citations.

Elicit's profile picture
Elicit1 month ago

Capability 5: Knowledge that compounds over time Improve each decision with Projects that keeps sources, context, instructions, and outputs connected in a single knowledge base.

Elicit's profile picture
Elicit1 month ago

One of the most important new things we did was create BioDecisionBench, a benchmark specifically evaluating LLMs on reasoning failures in high-stakes pharma decisions. BioDecisionBench tests for reasoning around preclinical study design, observational study design (e.g. target trial validity), and rigorous reasoning throughout the drug development process. When we evaluate on BioDecisionBench, Elicit highlights more key considerations for decisions (76.7% coverage in Elicit Smartest mode vs. 68.8% for Claude Opus 5 Max). Read the full evaluation blog:

Elicit's profile picture
Elicit1 month ago

Early customers are already using Research Agent to prioritize indications, shape clinical strategy, design trials, develop laboratory methods, inform market entry, and strengthen patent strategy. Read the launch post and see examples of Research Agent in action:

Related Videos

We’re entering the 10x speed of research publication workflow with AI. SciSpace (SciSpace), the first AI Agent built exclusively for the scientific community, is releasing so many inredibly useful features. 🎯 This is the AI Agent that can use 150+ tools, 59 databases, and 280M+ papers A few weeks back they launched BioMed Agent - It can design entire molecular biology workflows and even create publication-ready illustrations in a single prompt. This is its new domain-specialized AI co-scientist that sits on top of the existing SciSpace Agent and automates full biomedical workflows, from raw data and papers to analysis, decisions, and the final production-grade illustrations. You just need to give it 1 prompt. And today the added the following - Library Search, so it can search and analyze the PDFs already sitting in My Library, letting people ask questions across their own paper pile while keeping it private. - Now connects directly to Zotero, so the Agent can pull and work with the papers you already saved there without manual uploads. - For bigger prompts, it auto-triggers a Report Writing Sub-Agent that turns the chat into a structured research-style report, which is way cleaner for literature reviews and long summaries. - And when you get something worth keeping, Save to Notebook lets you store the output as .md notes with citations in My notebooks, so the work becomes reusable research notes instead of disappearing into chat. Behind the scenes, it indexes the PDF text, pulls a few relevant chunks for the question, then writes an answer grounded on those chunks.

Rohan Paul

11,574 views • 7 months ago