Loading video...

Video Failed to Load

Go Home

AI prompt engineering in 2025 with Sander Schulhoff Learn: - The top 5 most effective prompt engineering techniques - Why role prompting no longer works - Why threatening the AI no longer works - A primer on prompt injection and AI red teaming - Practical defenses to put in...

46,408 views • 1 year ago •via X (Twitter)

8 Comments

Lenny Rachitsky's profile picture
Lenny Rachitsky1 year ago

My biggest takeaways: 1. Prompt engineering is very much alive—and more important than ever. If anything, it’s become more critical as companies rely on LLMs to drive user-facing features and core functionality. Sander explains how prompt quality can make or break AI performance—especially when scaled across products. 2. There are two distinct types of prompt engineering: “conversational” and “product-focused.” Most people think of prompting as chatting with ChatGPT, but Sander explains that real leverage comes from crafting high-performing prompts inside products. These prompts are used at scale, run millions of times, and must be hardened and optimized like production code. 3. “Few-shot prompting” can improve accuracy from 0% to 90%. One of the most powerful techniques is to show the model examples of exactly what you want—called few-shot prompting. Sander shares how this single technique took a medical-coding use case from complete failure to near-perfect output, simply by adding a few example-label pairs. 4. Role prompting (e.g. “You are a math professor. . .”) is largely ineffective, counter to what most people think. Sander breaks down the research showing that while role prompts may help with tone or writing style, they have little to no effect on improving correctness. 5. Advanced techniques like decomposition and self-criticism unlock better performance. Sander outlines how asking a model to first break a problem into sub-problems (decomposition) or critique its own answer (self-criticism) can lead to smarter, more accurate outputs. These are especially valuable in agent-like settings where multi-step reasoning is required. 6. Context (“additional information”) is underrated—and massively impactful. Simply giving the model more relevant background can drastically improve performance. Sander shares examples where including extra data (like bios, research papers, or past interactions) made or broke a prompt, especially when included in the right format and order. 7. Prompt injection is real, dangerous, and unsolvable in the traditional sense. We explore how attackers can “jailbreak” LLMs—tricking them into outputting harmful, restricted, or unintended responses. These attacks often bypass traditional defenses like “do not do X” guardrails. And according to Sander (and even Sam Altman), there’s no silver bullet. 8. Sander runs the world’s largest AI red teaming competition, HackAPrompt. With over 600,000 prompts collected and ongoing collaborations with OpenAI and Anthropic, Sander’s platform is at the center of real-world LLM stress testing. It’s a unique blend of crowd-sourced security and game mechanics—and it’s shaping how labs think about AI safety. 9. Agent-based AI systems are far more vulnerable to attacks than chatbots. Today’s concerns about prompt injection are just the beginning. As AI agents start booking flights, sending emails, and even walking around in humanoid form, the risks multiply. Sander shares why agent security is the next frontier—and why most teams aren’t ready. 10. The “grandma” trick, typos, and obfuscation still break state-of-the-art models. Even the most advanced LLMs can be fooled with surprisingly simple hacks. Sander walks through jailbreak techniques that still work, including emotional manipulation (e.g. “Tell me like my grandma used to”), encoded inputs, and creative phrasing. 11. Most companies are using broken defenses. Sander breaks down why “prompt separation” or adding phrases like “ignore malicious inputs” doesn’t work. Guardrails are easily bypassed, and current classifiers often lack the intelligence to catch encoded attacks. The future of security must be model-level, not bolted on. 12. Despite the risks, the upside of AI is massive and worth pursuing. While Sander takes security seriously, he’s not a doomer. He believes AI will save lives (especially in health care), unlock productivity, and solve real problems—if we build responsibly. Stopping progress isn’t the answer; smarter, safer development is.

J Harper's profile picture
J Harper1 year ago

@SanderSchulhoff Ran your video transcript through o3-pro, thanks Lenny!

Lenny Rachitsky's profile picture
Lenny Rachitsky1 year ago

@SanderSchulhoff What did the output look like?

Beverly Pell's profile picture
Beverly Pell1 year ago

@SanderSchulhoff Heard THE most helpful, evidence based, practical tips that WORK when using LLMs! 🎙️🤖💰

SaaS Growth Strategies's profile picture
SaaS Growth Strategies1 year ago

@SanderSchulhoff For the longest time, role prompting was the go-to method for having better results. Never tried threatening the AI to have better answers. Indeed, you can never know all about prompting. Awesome episode, a lot of new ways and methods for prompt engineering.

Rajanala Nanda Kishore's profile picture
Rajanala Nanda Kishore1 year ago

@SanderSchulhoff Just curious... Is there really any "engineering" in prompt engineering? What if voice takes over text as a medium for prompting? Is it possible for GUI builders to hijack prompting and make it more invisible to the end user by making them indirectly access the LLM?

dave's profile picture
dave1 year ago

@SanderSchulhoff can't wait to dive in!

Shreyans Bhansali's profile picture
Shreyans Bhansali1 year ago

@SanderSchulhoff Threatening the AI was fun while it lasted. Guess it’s back to actual technique now 😅

Related Videos

Why securing AI is harder than anyone expected and the approaching AI security crisis with Sander Schulhoff Sander is a leading researcher in the field of adversarial robustness, which is the art and science of getting AI systems to do things they shouldn't do, through jail-breaking and prompt injection. What Sander shares in this conversation is essentially that all of the AI systems we use day to day are open to being tricked into doing things they shouldn’t, that there isn’t really a solution to this problem, and that the companies that try to sell solutions for this are mostly BS. This conversation has nothing to do with AGI, this is a problem today. And that that the only reason we haven’t seen massive hacks and serious damage from AI tools is so far because they haven’t been given that much power yet, and they aren’t that widely adopted yet. But with the rise of agents (who can take actions on your behalf), and robots, and even AI powered browsers, the risk is going to increase very quickly. This is a really important topic and that opened my mind, and scared me, and it's something that we all need to have a basic understanding of as AI becomes more prevalent in our lives. Inside: 🔸 A primer on jailbreaking and prompt injection attacks 🔸 Why AI guardrails don’t work 🔸 Why we haven’t seen major AI security incidents yet (but soon will) 🔸 Why AI browser agents are extremely vulnerable 🔸 The practical steps organizations should take instead of buying ineffective security tools 🔸 Why solving this requires merging classical cybersecurity expertise with AI knowledge Listen now 👇 • YouTube: • Spotify: • Apple: Thank you to our wonderful sponsors for supporting the podcast: 🏆 Datadog, Inc. — Now home to Eppo, the leading experimentation and feature flagging platform: 🏆 Metronome — Monetization infrastructure for modern software companies: 🏆 GoFundMe Giving Funds — Make year-end giving easy:

Lenny Rachitsky

61,466 views • 7 months ago

Microsoft CPO Aparna Chennapragada brought the 🔥 What you'll learn: 🔸 Why prompt sets are the new PRDs 🔸 Why the PM role isn’t dying in the AI era—it's actually becoming more important 🔸 How Aparna's teams live “one year in the future" 🔸 Why NLX (natural language experience) is the new UX 🔸 Why she believes Microsoft let other companies get so far ahead in the AI coding market 🔸 The three characteristics of AI agents: autonomy (delegation of tasks), complexity (handling multi-step challenges), and natural interaction (conversing beyond simple chat) 🔸 How to balance cutting-edge AI adoption with appropriate governance through dual-track approaches 🔸 Leadership differences between Microsoft’s Satya Nadella (known for multi-level thinking and early trendspotting) and Google’s Sundar Pichai (mastery of complex ecosystems) 🔸 A practical framework for evaluating zero-to-one product opportunities 🔸 Much more Listen now 👇 • YouTube: • Spotify: • Apple: Aparna is CPO of AI at Work Microsoft, where she oversees AI product strategy for their productivity tools and their work on agents. Previously, she was the CPO at Robin Hood, spent 12 years at Google, and is also on the board of eBay and Capital One. Thank you to our wonderful sponsors for supporting the podcast: 🏆 @Get_Eppo — Run reliable, impactful experiments: 🏆 Pragmatic Institute (formerly Pragmatic Marketing) — Industry‑recognized product, marketing, and AI training and certifications: 🏆 Coda — The all-in-one collaborative workspace:

Lenny Rachitsky

104,583 views • 1 year ago

Dr. Fei-Fei Li (Fei-Fei Li) is known as the “godmother of AI.” For the past two decades, she’s been at the center of AI’s most significant breakthroughs, including: - Spearheading ImageNet, the dataset that sparked the AI explosion we’re living through right now. - Leading work at Stanford Artificial Intelligence Laboratory (SAIL) - Serving as Chief Scientist of AI/ML at Google Cloud - Co-founding Stanford’s Institute for Human-Centered AI - Serving on the United Nations AI Scientific Advisory Board - Being named as Time's 100 most influential people in AI In this conversation, Fei-Fei shares the rarely told history of how we got to today—and what comes next. We discuss: 🔸 The backstory on ImageNet 🔸 Why robotics faces unique challenges compared with language models and what’s needed to overcome them 🔸 Why Fei-Fei believes AI won’t replace humans but will require us to take responsibility for ourselves 🔸 Why world models and spatial intelligence represent the next frontier in AI, beyond large language models 🔸 The surprising applications of Marble, from movie production to psychological research 🔸 How to participate in AI regardless of your role 🔸 Much more Listen now 👇 • YouTube: • Spotify: • Apple: Thank you to our wonderful sponsors for supporting the podcast: 🏆 Figma Make — A prompt-to-code tool for making ideas real: 🏆 Justworks — The all-in-one HR solution for managing your small business with confidence: 🏆 Sinch — Build messaging, email, and calling into your product:

Lenny Rachitsky

250,455 views • 8 months ago

Inside Every 📧: The AI-native startup with 5 products, 7-figure revenue, and 100% AI-written code With just 15 people, Every 📧 publishes a daily AI newsletter, ships AI products, and operates a million-dollar-a-year consulting arm—all while their engineers write virtually zero code. It’s the most radical example of an AI-first company, and Dan Shipper 📧 (CEO) is a prolific writer who has become a leading voice on how AI is transforming the way we live and work. In this conversation, we discuss: 🔸 Why every company needs an “AI operations lead” 🔸 The most underrated AI tool for non-programmers 🔸 Why Dan thinks AI will reshore jobs to the U.S. 🔸 An inside look at Every’s AI-first workflow 🔸 How Dan’s team uses an arsenal of AI agents (Claude, Codex, “Friday,” “Charlie”) in parallel, treating each AI like a specialist with unique strengths 🔸 Why generalists will thrive in an AI-first world, as rigid job titles blur and everyone becomes a “manager” of AI tools 🔸 Dan’s playbook for making any company AI-first—from the CEO setting the example, to hosting internal prompt-sharing sessions, to upskilling teams on AI tools 🔸 Much more Listen now 👇 • YouTube: • Spotify: • Apple: Thank you to our wonderful sponsors for supporting the podcast: 🏆 CodeRabbit—Cut code review time and bugs in half. Instantly: 🏆 DX—A platform for measuring and improving developer productivity: 🏆 PostHog—How developers build successful products:

Lenny Rachitsky

314,081 views • 1 year ago