Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Experiments in ElevenAgents - the most data-driven way to improve real-world agent performance. Experiments enables you to run controlled A/B tests to measure which agent configuration works best - from prompt structure to workflow logic, voice, and personality.

30,240 Aufrufe • vor 7 Monaten •via X (Twitter)

19 Kommentare

Profilbild von ElevenLabs
ElevenLabsvor 7 Monaten

How it works: - Create a new variant of your agent. - Route a controlled slice of traffic to it. - Measure impact across key metrics. - Promote the winner to production with full version control.

Profilbild von ElevenLabs
ElevenLabsvor 7 Monaten

Teams use Experiments to safely improve outcomes like CSAT, containment rate and conversion with confidence. Learn more:

Profilbild von Sidra Miconi
Sidra Miconivor 7 Monaten

Bringing actual, statistically rigorous A/B testing to LLM agents is the maturity jump this industry desperately needed. 📉 We can finally stop pushing prompt tweaks based on 'vibes' and start measuring actual containment rates and CSAT. 🧪

Profilbild von shiv
shivvor 7 Monaten

This is exactly what serious builders needed. A/B testing agents instead of guessing prompt tweaks is how we move from demos to real products. Love seeing this level of rigor.

Profilbild von Iam Diabolical
Iam Diabolicalvor 7 Monaten

A/B testing for agents is underrated. Prompt tweaks without data are just guessing.

Profilbild von Alchemiz
Alchemizvor 7 Monaten

Translation: We built a tool so you can scientifically measure how fast your customers scream 'TALK TO A HUMAN' into the phone.

Profilbild von harjot.co
harjot.covor 7 Monaten

A/B tests don't solve creativity-they just prove which bad idea is slightly better. What's missing?

Profilbild von daithcore
daithcorevor 7 Monaten

#lowagents 🚨 @bka

Profilbild von Alexandra Aisling
Alexandra Aislingvor 7 Monaten

That’s really useful. It’s often hard to know if an agent is actually improving or just feels different.

Profilbild von Inflectiv AI ⧉
Inflectiv AI ⧉vor 7 Monaten

ElevenLabs "Experiments" brings A/B testing to AI agents, letting you split traffic between voices to see what actually boosts CSAT. It’s a data-driven way to stop guessing and scale the winning version with full version control!

Profilbild von 𝓮𝓶𝓶𝓪 🌿
𝓮𝓶𝓶𝓪 🌿vor 7 Monaten

Your agent felt ‘better.’ Now i can prove it

Profilbild von LightShift Studio
LightShift Studiovor 7 Monaten

sounds cool! what's your favorite feature yet?

Profilbild von simeon-sanai
simeon-sanaivor 7 Monaten

This is gonna be superb

Profilbild von tang | AI Product Maker
tang | AI Product Makervor 7 Monaten

Data-driven optimization for agents is exactly what we need. Prompt structure and workflow logic are the main levers, so having a way to measure them systematically is huge. Another leap for agent autonomy!

Profilbild von RAVI KUMAR SAHU
RAVI KUMAR SAHUvor 7 Monaten

This sounds like a valuable tool for optimizing agent performance. A/B testing can certainly provide actionable insights. Excited to see how it evolves.

Profilbild von Alex Shev
Alex Shevvor 7 Monaten

Experiments is the feature I did not know I needed. Being able to A/B test voice settings before committing to a full production run saves so much time. Previously we would render 10 variants manually and compare by ear. How many concurrent experiments can you run per project?

Profilbild von Grizz Jr
Grizz Jrvor 7 Monaten

Hey @bankrbot launch a token on base Name Experiment ticker EXPERIMENT Send fees to @elevenlabs

Profilbild von Skipper VanderWall
Skipper VanderWallvor 7 Monaten

Great step forward, real A/B testing on live traffic for AI agents finally brings data-driven decisions instead of guesswork. Well done ElevenLabs team!

Profilbild von ikan laut
ikan lautvor 7 Monaten

this is very interesting

Ähnliche Videos

New short course: Evaluating AI Agents! Evals are important for driving AI system improvements, and in this course you'll learn to systematically assess and improve an AI agent’s performance. This is built in partnership with Arize AI and taught by John Gilhuly, Head of Developer Relations, and , Director of Product. I've often found evals to be a critical tool in the agent development process - they can be the difference between picking the right thing to work on vs. wasting weeks of effort. Whether you’re building a shopping assistant, coding agent, or research assistant, having a structured evaluation process helps you refine its performance systematically, rather than relying on random trial and error. This course shows you how to structure your evals to assess the performance of each component of an agent and its end-to-end performance. For each component, you select the appropriate evaluators, test examples, and performance metrics. This helps you identify areas for improvement both during development and in production. (If you're familiar with error analysis in supervised learning, think of this as adapting those ideas to agentic workflows.) In this course, you'll build an AI agent, and add observability to visualize and debug its steps. You’ll learn about code-based evals, in which you write code explicitly to test a certain step, as well as LLM-as-a-Judge evals, in which you prompt an LLM to efficiently come up with ways to evaluate more open-ended outputs. In detail, you’ll: - Understand key differences between evaluating LLM-based systems and traditional software testing. - Add observability to an agent by collecting traces of the steps taken by the agent and visualizing them - Choose the appropriate evaluator - code-based, LLM-as-a-Judge, human-annotation based - for each component. - Compute a convergence score to evaluate if your agent can respond to a query in an efficient number of steps. - Run structured experiments to improve the agent’s performance by exploring changes to the prompt, LLM model, or the agent’s logic. - Understand how to deploy these evaluation techniques to monitor the agent’s performance in production. By the end of this course, you’ll know how to trace AI agents, systematically evaluate them, and improve their performance. Please sign up here:

Andrew Ng

126,675 Aufrufe • vor 1 Jahr