正在加载视频...
视频加载失败
Your users never use your product the way you designed it. That’s why we’re launching Prompt Simulations today. Pick a prompt → generate realistic users and scenarios → run full multi-turn conversations → see exactly where your prompt breaks. Test against real behavior before real users find the edge cases.
210,543 次观看 • 10 天前 •via X (Twitter)
21 条评论

How it works: 1. Pick a saved prompt 2. Generate up to 25 realistic scenarios with different personas, goals, and pass/fail criteria 3. Run full multi-turn conversations automatically 4. See which conversations passed, failed, or were inconclusive Try Prompt Simulations →

multi-turn is where a lot of the interesting failures start showing up

This is much closer to testing actual product behavior instead of just testing individual responses

this has been one of the biggest gaps for us in prompt testing. the prompt can look perfect in isolation and still completely fall apart once a real conversation starts

love the idea of generating different personas instead of testing the same clean input over and over

Love the focus on real user behavior, not just ideal inputs

this is a much better way to test prompts.

This is a great way to find hidden issues.

a prompt can look completely fine for the first few turns and then suddenly go off the rails. this catches that nicely

being able to rerun the same scenarios after changing the prompt is probably my favorite part

users will always find a way to take the conversation somewhere you did not expect lol

@RespanAI oh man, this is so true. built an AI tool myself and had to constantly adapt because of unexpected user behavior. simulations can be a game changer!

The multi-turn testing part is especially interesting to me.

really like that the scenarios are reusable across prompt versions. makes iteration much easier

testing prompts with 25 different scenarios before real users hit the edge cases is a very practical workflow.

This makes testing actually useful

seeing why something failed instead of just getting a pass or fail is really nice

Prompt simulations get much more useful if each failed run leaves a breakage receipt: prompt version, simulated user profile, turn where intent drifted, violated constraint, and patch tried. Otherwise the demo only proves the happy user was polite.

good approach,testing realistic user behavior before launch gives teams a clearer view of where prompts fail

Cool

Testing against real behavior makes prompts far more resilient.
