Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing Experiments in ElevenAgents - the most data-driven way to improve real-world agent performance. Experiments enables you to run controlled A/B tests to measure which agent configuration works best - from prompt structure to workflow logic, voice, and personality.

30,240 görüntüleme • 7 ay önce •via X (Twitter)

19 Yorum

ElevenLabs profil fotoğrafı
ElevenLabs7 ay önce

How it works: - Create a new variant of your agent. - Route a controlled slice of traffic to it. - Measure impact across key metrics. - Promote the winner to production with full version control.

ElevenLabs profil fotoğrafı
ElevenLabs7 ay önce

Teams use Experiments to safely improve outcomes like CSAT, containment rate and conversion with confidence. Learn more:

Sidra Miconi profil fotoğrafı
Sidra Miconi7 ay önce

Bringing actual, statistically rigorous A/B testing to LLM agents is the maturity jump this industry desperately needed. 📉 We can finally stop pushing prompt tweaks based on 'vibes' and start measuring actual containment rates and CSAT. 🧪

shiv profil fotoğrafı
shiv7 ay önce

This is exactly what serious builders needed. A/B testing agents instead of guessing prompt tweaks is how we move from demos to real products. Love seeing this level of rigor.

Iam Diabolical profil fotoğrafı
Iam Diabolical7 ay önce

A/B testing for agents is underrated. Prompt tweaks without data are just guessing.

Alchemiz profil fotoğrafı
Alchemiz7 ay önce

Translation: We built a tool so you can scientifically measure how fast your customers scream 'TALK TO A HUMAN' into the phone.

harjot.co profil fotoğrafı
harjot.co7 ay önce

A/B tests don't solve creativity-they just prove which bad idea is slightly better. What's missing?

daithcore profil fotoğrafı
daithcore7 ay önce

#lowagents 🚨 @bka

Alexandra Aisling profil fotoğrafı
Alexandra Aisling7 ay önce

That’s really useful. It’s often hard to know if an agent is actually improving or just feels different.

Inflectiv AI ⧉ profil fotoğrafı
Inflectiv AI ⧉7 ay önce

ElevenLabs "Experiments" brings A/B testing to AI agents, letting you split traffic between voices to see what actually boosts CSAT. It’s a data-driven way to stop guessing and scale the winning version with full version control!

𝓮𝓶𝓶𝓪 🌿 profil fotoğrafı
𝓮𝓶𝓶𝓪 🌿7 ay önce

Your agent felt ‘better.’ Now i can prove it

LightShift Studio profil fotoğrafı
LightShift Studio7 ay önce

sounds cool! what's your favorite feature yet?

simeon-sanai profil fotoğrafı
simeon-sanai7 ay önce

This is gonna be superb

tang | AI Product Maker profil fotoğrafı
tang | AI Product Maker7 ay önce

Data-driven optimization for agents is exactly what we need. Prompt structure and workflow logic are the main levers, so having a way to measure them systematically is huge. Another leap for agent autonomy!

RAVI KUMAR SAHU profil fotoğrafı
RAVI KUMAR SAHU7 ay önce

This sounds like a valuable tool for optimizing agent performance. A/B testing can certainly provide actionable insights. Excited to see how it evolves.

Alex Shev profil fotoğrafı
Alex Shev7 ay önce

Experiments is the feature I did not know I needed. Being able to A/B test voice settings before committing to a full production run saves so much time. Previously we would render 10 variants manually and compare by ear. How many concurrent experiments can you run per project?

Grizz Jr profil fotoğrafı
Grizz Jr7 ay önce

Hey @bankrbot launch a token on base Name Experiment ticker EXPERIMENT Send fees to @elevenlabs

Skipper VanderWall profil fotoğrafı
Skipper VanderWall7 ay önce

Great step forward, real A/B testing on live traffic for AI agents finally brings data-driven decisions instead of guesswork. Well done ElevenLabs team!

ikan laut profil fotoğrafı
ikan laut7 ay önce

this is very interesting

Benzer Videolar

New short course: Evaluating AI Agents! Evals are important for driving AI system improvements, and in this course you'll learn to systematically assess and improve an AI agent’s performance. This is built in partnership with Arize AI and taught by John Gilhuly, Head of Developer Relations, and , Director of Product. I've often found evals to be a critical tool in the agent development process - they can be the difference between picking the right thing to work on vs. wasting weeks of effort. Whether you’re building a shopping assistant, coding agent, or research assistant, having a structured evaluation process helps you refine its performance systematically, rather than relying on random trial and error. This course shows you how to structure your evals to assess the performance of each component of an agent and its end-to-end performance. For each component, you select the appropriate evaluators, test examples, and performance metrics. This helps you identify areas for improvement both during development and in production. (If you're familiar with error analysis in supervised learning, think of this as adapting those ideas to agentic workflows.) In this course, you'll build an AI agent, and add observability to visualize and debug its steps. You’ll learn about code-based evals, in which you write code explicitly to test a certain step, as well as LLM-as-a-Judge evals, in which you prompt an LLM to efficiently come up with ways to evaluate more open-ended outputs. In detail, you’ll: - Understand key differences between evaluating LLM-based systems and traditional software testing. - Add observability to an agent by collecting traces of the steps taken by the agent and visualizing them - Choose the appropriate evaluator - code-based, LLM-as-a-Judge, human-annotation based - for each component. - Compute a convergence score to evaluate if your agent can respond to a query in an efficient number of steps. - Run structured experiments to improve the agent’s performance by exploring changes to the prompt, LLM model, or the agent’s logic. - Understand how to deploy these evaluation techniques to monitor the agent’s performance in production. By the end of this course, you’ll know how to trace AI agents, systematically evaluate them, and improve their performance. Please sign up here:

Andrew Ng

126,675 görüntüleme • 1 yıl önce