Loading video...

Video Failed to Load

Go Home

“We purposely build or discover situations where models might be behaving in misaligned ways” Evan Hubinger discusses stress-testing AI by creating “model organisms” to study failure points and refine model safeguards under Anthropic's Responsible Scaling Policy.

1,643,782 views • 1 year ago •via X (Twitter)

9 Comments

FAR.AI's profile picture
FAR.AI1 year ago

@AnthropicAI Follow us for AI safety insights And watch the full video

SujyotChapade12's profile picture
SujyotChapade121 year ago

@EvanHub @AnthropicAI Exciting innovation from @EvanHub ! The "model organisms" strategy for stress-testing AI uncovers vulnerabilities and strengthens safeguards, ensuring #AI evolves safely and responsibly.

Ben Schulz's profile picture
Ben Schulz1 year ago

@EvanHub @AnthropicAI This is kinda promising. Great idea.

Rohit Prasad ✨'s profile picture
Rohit Prasad ✨1 year ago

Great insight from @EvanHub on stress-testing AI! By purposely creating "model organisms" to explore misaligned behaviors, we can identify failure points and refine safeguards. This approach under @AnthropicAI's Responsible Scaling Policy is a crucial step toward building safer, more reliable AI systems. It's all about learning from potential risks before they become real issues.

Naeem ul Fateh's profile picture
Naeem ul Fateh1 year ago

@EvanHub @AnthropicAI A model must need to be reevaluated before a commission before that should be put it into academic or industrial research to have effective worthy outcome.

KI-CREW's profile picture
KI-CREW1 year ago

@EvanHub @AnthropicAI those are the real guys

Dirac's profile picture
Dirac1 year ago

@EvanHub @AnthropicAI @grok Hi

Ted Licious 🦭's profile picture
Ted Licious 🦭1 year ago

@EvanHub @AnthropicAI So AI gain of function research? Sounds… safe?

𝔗𝔥𝔢𝔬 𝔊𝔬𝔱𝔱𝔴𝔞𝔩𝔡's profile picture
𝔗𝔥𝔢𝔬 𝔊𝔬𝔱𝔱𝔴𝔞𝔩𝔡1 year ago

@EvanHub @AnthropicAI "Alignment" is just another word for "censorship". You will waste a lot oft time and we will use chinese alternatives that are not "aligned".

Related Videos

Responsible AI should not feel like the path of most resistance. That was the strongest idea from my SAS Innovate conversation with Reggie Townsend, leading Data Ethics, Governance, and Social Impact at SAS Software. Reggie framed governance not as a compliance layer, but as a way to scale human judgment. Too often, we innovate first and govern later. Model selected. Agent deployed. Process built. Then governance arrives at the end and feels like friction. Bolt it on after the fact, and resistance is guaranteed. The opportunity: design responsible AI, so it becomes intuitive, action-oriented, and useful in the flow of work. That is what Reggie meant by making responsibility "irresistible." Second point: Use cases must lead. When everyone can access the same models, differentiation will not come from the technology. It will come from how leaders define outcomes, govern applications, and connect business value to institutional trust. The risk does not live in the model. The value does not live in the model. Both live in the use case. For CEOs and boards, this is the shift: from model-first oversight to outcome-first accountability. Better questions before scaling AI: - What human decision are we shaping? - What business outcome are we improving? - What risk are we containing? - What judgment are we extending? Responsible AI becomes strategic when it helps people make better decisions, faster, with greater confidence. Most leaders can't see where governance sits inside their AI operating model. SAS AI Navigator makes it visible: Design for the human. Not only the technology.

Sabine VanderLinden

835,838 views • 3 months ago

After taking some time off post-Rapid, I'm excited to share what I’ve been up to since: Datawizz AI! We’ve raised a $12.5M Seed led by Human Capital to make AI 10x cheaper, 2x more accurate and 15x faster by transitioning from LLMs to SLMs. AI is eating the world. But unit economics are eating AI. Looking at the fastest growing AI products, they all share two traits - growing fast, and painful inference bills. General-purpose LLMs are just too expensive to run. A big reason for that is we train LLMs to be good at everything - answer any question, be an expert on any topic. The big labs dub this "generalisation", but for real-world applications, it is unnecessary. In reality - many AI applications need models to be experts in one thing - and do that thing extremely well. Your coding model doesn’t need to memorize ancient recipes for Garum sauce. This is where Datawizz comes in - we sit between the AI applications and automatically create smaller (100x-1,000x) specialized models to handle specific aspects of your work. By focusing the model and combining industry-data in the distillation process - we end up with models that beat SOTA LLMs at a fraction of the cost. We created Datawizz to make AI specialized and scalable. We’re early in the journey, but have already been able to save companies 90%+ on their inference bill and speed up their apps by 10x. Excited to build better AI platforms? Join the Datawizz team (link in first comment)

Iddo Gino 🐙

21,928 views • 11 months ago

Today, we’re launching Parsed. We are incredibly lucky to live in a world where we stand on the shoulders of giants, first in science and now in AI. Our heroes have gotten us to this point, where we have brilliant general intelligence in our pocket. But this is a local minima. We now have an ecosystem of burgeoning tasks where each requires a different kind of intelligence, a different context, a whole host of implicit assumptions and latent knowledge and domain expertise that is very difficult to cram into a system prompt. The big labs want you renting their $50k/month amnesiac interns that forget everything between conversations. Generic behemoths that get quantised, versioned and deprecated behind the scenes, where the only element of control you have is your messy monolithic user prompt. We want people who need their own intelligence to be able to not only access it, but also control it. And whilst the big general models are unbelievably good chatbots and coding agents and purveyors of the world, specialisation of intelligence is required. Clinical scribes, marketing compliance agents, legal red-lining models, insurance policy recommenders, the list goes on. And so that’s what Parsed does: deploy your own frontier model that actually learns. We eval your specific task, build a custom evaluation harness, optimise a model just for you, and host it with continual learning. We bake all the context and knowledge of your task into the model itself, from your engineers to your domain experts to customer feedback, all in a tight SFT → RL loop, with useful interpretability made possible by the open-source ecosystem we build on top of. No more 2000-word prompts with seventeen "IMPORTANT: NEVER DO X" clauses. Your model gets better at YOUR job every single day; the amnesiac pseudo-gods have had their run. Your model, your data, your moat. Let's build 🫡

Charlie O'Neill

137,812 views • 1 year ago