Loading video...

Video Failed to Load

Go Home

Introducing Experiments in ElevenAgents - the most data-driven way to improve real-world agent performance. Experiments enables you to run controlled A/B tests to measure which agent configuration works best - from prompt structure to workflow logic, voice, and personality.

30,240 views • 7 months ago •via X (Twitter)

19 Comments

ElevenLabs's profile picture
ElevenLabs7 months ago

How it works: - Create a new variant of your agent. - Route a controlled slice of traffic to it. - Measure impact across key metrics. - Promote the winner to production with full version control.

ElevenLabs's profile picture
ElevenLabs7 months ago

Teams use Experiments to safely improve outcomes like CSAT, containment rate and conversion with confidence. Learn more:

Sidra Miconi's profile picture
Sidra Miconi7 months ago

Bringing actual, statistically rigorous A/B testing to LLM agents is the maturity jump this industry desperately needed. 📉 We can finally stop pushing prompt tweaks based on 'vibes' and start measuring actual containment rates and CSAT. 🧪

shiv's profile picture
shiv7 months ago

This is exactly what serious builders needed. A/B testing agents instead of guessing prompt tweaks is how we move from demos to real products. Love seeing this level of rigor.

Iam Diabolical's profile picture
Iam Diabolical7 months ago

A/B testing for agents is underrated. Prompt tweaks without data are just guessing.

Alchemiz's profile picture
Alchemiz7 months ago

Translation: We built a tool so you can scientifically measure how fast your customers scream 'TALK TO A HUMAN' into the phone.

harjot.co's profile picture
harjot.co7 months ago

A/B tests don't solve creativity-they just prove which bad idea is slightly better. What's missing?

daithcore's profile picture
daithcore7 months ago

#lowagents 🚨 @bka

Alexandra Aisling's profile picture
Alexandra Aisling7 months ago

That’s really useful. It’s often hard to know if an agent is actually improving or just feels different.

Inflectiv AI ⧉'s profile picture
Inflectiv AI ⧉7 months ago

ElevenLabs "Experiments" brings A/B testing to AI agents, letting you split traffic between voices to see what actually boosts CSAT. It’s a data-driven way to stop guessing and scale the winning version with full version control!

𝓮𝓶𝓶𝓪 🌿's profile picture
𝓮𝓶𝓶𝓪 🌿7 months ago

Your agent felt ‘better.’ Now i can prove it

LightShift Studio's profile picture
LightShift Studio7 months ago

sounds cool! what's your favorite feature yet?

simeon-sanai's profile picture
simeon-sanai7 months ago

This is gonna be superb

tang | AI Product Maker's profile picture
tang | AI Product Maker7 months ago

Data-driven optimization for agents is exactly what we need. Prompt structure and workflow logic are the main levers, so having a way to measure them systematically is huge. Another leap for agent autonomy!

RAVI KUMAR SAHU's profile picture
RAVI KUMAR SAHU7 months ago

This sounds like a valuable tool for optimizing agent performance. A/B testing can certainly provide actionable insights. Excited to see how it evolves.

Alex Shev's profile picture
Alex Shev7 months ago

Experiments is the feature I did not know I needed. Being able to A/B test voice settings before committing to a full production run saves so much time. Previously we would render 10 variants manually and compare by ear. How many concurrent experiments can you run per project?

Grizz Jr's profile picture
Grizz Jr7 months ago

Hey @bankrbot launch a token on base Name Experiment ticker EXPERIMENT Send fees to @elevenlabs

Skipper VanderWall's profile picture
Skipper VanderWall7 months ago

Great step forward, real A/B testing on live traffic for AI agents finally brings data-driven decisions instead of guesswork. Well done ElevenLabs team!

ikan laut's profile picture
ikan laut7 months ago

this is very interesting

Related Videos

New short course: Evaluating AI Agents! Evals are important for driving AI system improvements, and in this course you'll learn to systematically assess and improve an AI agent’s performance. This is built in partnership with Arize AI and taught by John Gilhuly, Head of Developer Relations, and , Director of Product. I've often found evals to be a critical tool in the agent development process - they can be the difference between picking the right thing to work on vs. wasting weeks of effort. Whether you’re building a shopping assistant, coding agent, or research assistant, having a structured evaluation process helps you refine its performance systematically, rather than relying on random trial and error. This course shows you how to structure your evals to assess the performance of each component of an agent and its end-to-end performance. For each component, you select the appropriate evaluators, test examples, and performance metrics. This helps you identify areas for improvement both during development and in production. (If you're familiar with error analysis in supervised learning, think of this as adapting those ideas to agentic workflows.) In this course, you'll build an AI agent, and add observability to visualize and debug its steps. You’ll learn about code-based evals, in which you write code explicitly to test a certain step, as well as LLM-as-a-Judge evals, in which you prompt an LLM to efficiently come up with ways to evaluate more open-ended outputs. In detail, you’ll: - Understand key differences between evaluating LLM-based systems and traditional software testing. - Add observability to an agent by collecting traces of the steps taken by the agent and visualizing them - Choose the appropriate evaluator - code-based, LLM-as-a-Judge, human-annotation based - for each component. - Compute a convergence score to evaluate if your agent can respond to a query in an efficient number of steps. - Run structured experiments to improve the agent’s performance by exploring changes to the prompt, LLM model, or the agent’s logic. - Understand how to deploy these evaluation techniques to monitor the agent’s performance in production. By the end of this course, you’ll know how to trace AI agents, systematically evaluate them, and improve their performance. Please sign up here:

Andrew Ng

126,675 views • 1 year ago