Dan Shipper's banner
Dan Shipper's profile picture

Dan Shipper

@danshipper122,224 subscribers

ceo @every | the only subscription you need to stay at the edge of AI

Shorts

BREAKING: When today’s jobs are automated by AI, what will great human work look like? This is the most important question of our time. Introducing Thesis Statements, a new project from Every 🪨 bringing together 100 builders and thinkers to call their shot: We asked them to make a specific prediction about what great human work will look like after automation. Today we’re launching the first 25 Thesis Statements from an incredible group including: • Karri Saarinen • Chris Pedregal • Anne-Laure Le Cunff • Yash Tekriwal • Alex Komoroske • Tina He • Jonny Miller • Paul Millerd • sari azout • Tom Critchlow • Simone Stolzoff And 14 more amazing builders and thinkers. At Every 🪨 we believe there is a bright future for human work after automation. And we believe that there’s a small group of humans who know what it looks like—because they live the answers every day. But their ideas are still largely missing from the mainstream discourse about AI. That’s why we’re creating a public record of what people at the frontier are seeing now, so we can get these ideas to as many people as possible. We’ll also revisit them over time, and ask: Which claims held up? Which didn’t? Which became more useful as the technology changed—and which dissolved on contact with the world? Read them, argue with them, share them, and submit your own:

BREAKING: When today’s jobs are automated by AI, what will great human work look like? This is the most important question of our time. Introducing Thesis Statements, a new project from Every 🪨 bringing together 100 builders and thinkers to call their shot: We asked them to make a specific prediction about what great human work will look like after automation. Today we’re launching the first 25 Thesis Statements from an incredible group including: • Karri Saarinen • Chris Pedregal • Anne-Laure Le Cunff • Yash Tekriwal • Alex Komoroske • Tina He • Jonny Miller • Paul Millerd • sari azout • Tom Critchlow • Simone Stolzoff And 14 more amazing builders and thinkers. At Every 🪨 we believe there is a bright future for human work after automation. And we believe that there’s a small group of humans who know what it looks like—because they live the answers every day. But their ideas are still largely missing from the mainstream discourse about AI. That’s why we’re creating a public record of what people at the frontier are seeing now, so we can get these ideas to as many people as possible. We’ll also revisit them over time, and ask: Which claims held up? Which didn’t? Which became more useful as the technology changed—and which dissolved on contact with the world? Read them, argue with them, share them, and submit your own:

241,310 views

BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week Every 🪨 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 🪨:

BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week Every 🪨 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 🪨:

354,559 views

in software, we're entering a new age of skyscrapers: rather than a small group of humans hand-typing code into a computer, software is now built by thousands of humans and agents working together to build applications at a scale that was previously unthinkable. the unsung heroes of this new era are infrastructure engineers. they are the people who are figuring out how to make software with 100x the amount of code, and 1000x the amount of contributors actually work OpenAI is one of the few companies already operating at this scale so we Every 🪨 spent a few days deep inside of the company to learn how their infrastructure team thinks about building software skyscrapers—and how they use agents to do it one of the most important articles we've written this year:

in software, we're entering a new age of skyscrapers: rather than a small group of humans hand-typing code into a computer, software is now built by thousands of humans and agents working together to build applications at a scale that was previously unthinkable. the unsung heroes of this new era are infrastructure engineers. they are the people who are figuring out how to make software with 100x the amount of code, and 1000x the amount of contributors actually work OpenAI is one of the few companies already operating at this scale so we Every 🪨 spent a few days deep inside of the company to learn how their infrastructure team thinks about building software skyscrapers—and how they use agents to do it one of the most important articles we've written this year:

14,699 views

codex-native weekend hack project: 1. buy cable to connect MIDI keyboard to computer 2. "hey codex, make a watcher script and a little web app to show me which chords im playing" 3. okay cool, now give me some exercises and help me see how to improve! literally 5 minutes start to finish, and it works flawlessly

codex-native weekend hack project: 1. buy cable to connect MIDI keyboard to computer 2. "hey codex, make a watcher script and a little web app to show me which chords im playing" 3. okay cool, now give me some exercises and help me see how to improve! literally 5 minutes start to finish, and it works flawlessly

32,894 views

BREAKING: is your inbox a dumpster fire with 45,000 unreads? take it to 0 emails in 5 minutes—safely Declare inbox bankruptcy with Cora. it's reversible, smart, and totally free so you can start fresh with a fresh inbox. declare inbox bankruptcy today:

BREAKING: is your inbox a dumpster fire with 45,000 unreads? take it to 0 emails in 5 minutes—safely Declare inbox bankruptcy with Cora. it's reversible, smart, and totally free so you can start fresh with a fresh inbox. declare inbox bankruptcy today:

60,146 views

Videos

danshipper's profile picture

BREAKING: Anthropic just dropped Fable 5.1—and CLAUDE IS SO BACK. We’ve spent the last week testing it at Every 🪨 across coding, writing, and knowledge work. Our verdict: It's finally Fable for everyone. It’s the strongest coding model we’ve used, but now it's fast, token-efficient, and CRUCIALLY actually speaks like a normal person. Here’s our vibe check: - A monster at coding. Kieran Klaassen rebuilt a working version of Proof, our document editor, from one prompt. It added useful details he hadn’t requested, and it handles enormous coding jobs that run for days at a time. It built a computer use Mac app for me called Hands in one-shot that other models failed at. - A Claude our writers want to use again. It has clearer prose, fewer AI tells, and it takes an edit without arguing. It's a significant upgrade over Opus 5. And won Katie Parrott's heart back. - About half the tokens as Opus 5, and much faster. In our Slack-agent tests, it delivered comparable results to Opus 5 using about half as many tokens, in about 60% of the time. - Knowledge work you can delegate. It can produce great knowledge work—like slide decks—end to end without making slop. And flew threw hammer's tests with flying colors. - It now supports zero-data-retention agreements. Now businesses can actually use it! A big barrier to Fable adoption is gone. Net Result: It's obviously an Opus 5 killer. If that was your daily driver you should switch today. If you're using GPT-5.6 in ChatGPT for Work, it's spinning the wheels on for big delegated tasks. I still use ChatGPT for Work more day to day, but I use way more tokens in Fable 5.1. I send it off at the beginning of the day to do big programming projects, like end to end MVP builds, and check in every once in a while. State of Play: The big knock on Anthropic was they built a supergenius in a datacenter that was almost unusable. It was too slow, argued back, and talked in technical gibberish. They've managed to solve those problems and more with Fable 5.1!

Dan Shipper

198,224 views • 2 days ago

danshipper's profile picture

BREAKING: Claude Opus 5 is OUT NOW! And…it’s a hard model to love. We’ve spent the last week Every 🪨 testing it across coding, writing, knowledge work, and our internal agent. It argued with instructions, stopped before the work was finished, and generally didn’t play well with our existing skills and plugins like Compound Engineering. Our first reaction was: What have they done to my boy? Then we deleted our existing skills and started from scratch. Without the elaborate workflows we had built for earlier models, Opus 5 got dramatically better, and even showed flashes of brilliance. Here’s our Day 0 vibe check: - It’s a poor man’s Fable. It has many of Fable’s personality quirks without Fable’s genius. - It breaks backward compatibility. If you’re using it with existing skills and workflows, watch out. It will often stop early or otherwise miss your instructions. - If you start from scratch, you’ll have better results. Kieran Klaassen figured out that if he just started from scratch without his existing skills, he could get dramatically better results. This is a model that takes some time to rebuild your workflows around—but if you do, there’s a payoff waiting. - Medium or low effort works better. @KieranKlassenn also found better results using Opus 5 on lower thinking levels. It seems that the more time you give it to think, the more likely it is to do the more annoying behaviors. Don’t just switch to Sonnet for a faster response! Try low thinking. I have two slots in my workflow: 1. The genius model I use for my biggest hardest tasks, currently Fable. 2. The smart, fast generalist I use for everything else, currently GPT-5.6. Opus 5 has the personality of the genius, but doesn’t have its top end. So that puts it in a strange middle ground that doesn’t really have a home in my day to day. I think I’ll use it mostly when I run out of Fable tokens. full vibe check on Every 🪨 in the next tweet 👇

Dan Shipper 📧

754,329 views • 1 month ago

danshipper's profile picture

BREAKING: Introducing Thesis, Every 🪨’s annual conference dedicated to answering the most important question in AI: What does great human work look like after automation? We believe there is a small group of humans who already know the answer to this question because they’re living it every day. But they're scattered across companies and industries, with few opportunities to learn from each other. That’s why we’re throwing Thesis November 5th, 2026 in Brooklyn New York. Our first speakers include: - Ivan Zhao (Ivan Zhao)—Founder and CEO, Notion - Andrew Ambrosino (Andrew Ambrosino)—Member of technical staff, Codex, OpenAI - Cat de Jong - Head of applied AI, Anthropic - Nick Thompson (nxthompson)—CEO, the Atlantic - Josh Miller (Josh Miller)—CEO and cofounder, The Browser Company - Cristobal Valenzuela (Cristóbal Valenzuela)—Co-CEO and cofounder, Runway - Lauren Reeder (Lauren Reeder)—Partner, Sequoia - Sahil Lavingia (Sahil Lavingia)—Founder, Gumroad - Riley Brown (Riley Brown)—Cofounder, Vibecode - Natalie Fratto (Natalie Fratto)—Founder and creator, Charts & Crafts - Allie Garfinkle (Allie Garfinkle)—Senior writer and editor, Fortune - Kane Kallaway (Kallaway)—Founder, Wavy Labs - Nat Eliason (Nat Eliason)—Head of Founders School, Alpha School - Kate Lee (Kate Lee)—Editor in chief, Every - Katie Parrott (Katie Parrott)—Staff writer, Every - Kieran Klaassen (Kieran Klaassen)—General manager of Cora, Every (With more very special people to announce soon!) Learn more: What to expect Thesis will feature talks, demonstrations, working sessions, office hours, and small-group conversations with builders, operators, and execs who are already using AI to do incredible human work. And, of course, you can bring your agent. Why New York New York is where the AI wave hits the beach: It's where new model capabilities meet real-world work. Because of this, it’s the best place in the world to see what happens when frontier technology leaves the lab and enters everyday life. That’s why we’re holding Thesis at Pioneer Works, a cultural center in Red Hook dedicated to blending art, science, music, and technology. Space is limited, so we’re accepting attendees by application. We’ll also livestream it for free for anyone who can’t attend in person. You should apply below. Apply to Thesis:

Dan Shipper 📧

161,415 views • 21 days ago

danshipper's profile picture

BREAKING: Anthropic just dropped Claude Fable 5—this is Mythos, made safe for public release. It is the best coding model in the world. We've been testing it internally Every 🪨 for the last week or so across coding, writing, marketing, editing, and more—here's our vibe check: - It broke our benchmarks. Fable scored a 91/100 on our Senior Engineer benchmark—this is human senior engineer level. The previous high score was Opus 4.8 at 63. GPT-5.5 is a 62. - It's a one-shot wonder. You can set it and forget for hours or overnight on huge coding tasks, and come back to completed work. It cleared entire production bug backlogs, built a playable 3D, and even made a 2-minute animated film—all one-shot. - Taste and attention to detail. In coding and knowledge work tasks, it has much better taste and attention to detail than we've ever seen. It gets subtle things right, adds little features you might not have thought of, and generally understands the assignment in ways that surprised us. - Great use of context. We set it loose analyzing customer feedback surveys and our website data and it came back with a crisp, clean report that identified a. our biggest problem and b. a concrete testable solution—and then we sent it off to build that. - It's best for power users. If you're already used to orchestrating multiple agents in your work, this model can do things that you've never seen before. If you're a knowledge worker or vibe coder with a more basic setup, you're not going to notice a huge difference—in fact, it probably isn't the right model for you. - It's very slow, token-hungry. Using this thing for regular knowledge work is like squashing an ant with a rocket launcher. It also routinely uses 500k to 1M tokens on tasks. That's why it's best for your heaviest jobs—but not as good for tasks like collaborative writing. - It's expensive. It's about twice as expensive as Opus, and it's also incredibly token hungry—so expect it to be something you'll use sparingly unless your company pays for it. Overall, I think of it like a warp drive for coding: It can get you across the galaxy in a few hours, when it used to take months or years. But it's not appropriate for getting around town—you need something faster, cheaper, and more maneuverable. The ceiling is extraordinarily high on this model though. Even our most advanced testers like Kieran Klaassen felt like they were only scratching the surface of it. Want our full vibe check with all of our testing and benchmarks? Read it on Every 🪨:

Dan Shipper

621,681 views • 2 months ago

danshipper's profile picture

BREAKING: Introducing All Access from Every 🪨, our new membership tier for the best builders in AI All Access subs get the Builder Pack which includes $7,000 in credits and free usage to the models + tool stack we use Every 🪨. All Access subscribers get: - $1,000 in Codex / @ChatGPTapp for Work credits - 12 months free of Cursor Pro+ - $4,000 in PostHog credits including self-driving to automatically fix bugs and identify issues in your production app - 1 year free of Framer - 6 months free of Notion And much more! (Did I mention $1,000 in Codex credits? It's time to build!) Get all access: Why All Access and the Builder Pack This is the best time in history to build something. For a long time, it’s been possible to one-shot impressive demos, but they’d fall flat the minute they hit production. But the release of GPT-5.6-Sol and Fable 5 heralds a new era: Everyone can build, launch, and maintain the software that they’ve always dreamed of. Everyone is a builder now. There’s just one catch: Building with AI is very expensive. (Ask me how I know.) (Alright, I’ll tell you. I accidentally used 2 billion tokens overnight this week on a big GPT-5.6-Sol run. Worth it.) This is unique in the history of technology. For most of the personal computing era, a billionaire and a solo builder could buy essentially the same top-of-the-line Mac. AI changes that: The more tokens you can afford, the more you can make. And we want to make that accessible to more people. That’s why the main feature of our new All Access plan is the Builder Pack: more than $7,000 in credits and discounts on the full stack we use to run Every, from idea to production—Codex, Claude, PostHog, Render, Gemini, FLORA, and more. Early-bird membership is only $500/year for the next 24 hours—and the Codex credits alone are worth $1,000. (I could’ve used it for my overnight run this week.) Now we’re handing it to you. Get all access: Meet the Builder Pack It's got more than $7,000 in offers from 10 of the AI products we use to write, design, build, and run Every 🪨: BUILD - $1,000 in Codex credits plus one month of ChatGPT for business - Twelve months free of Cursor Pro+ - One month free of Claude Max - Three months free of Google AI Pro DESIGN - One year free of Framer Pro - One month free of FLORA © Max HOST - $300 in Render credits IMPROVE - $4,000 in PostHog credits - Six months free of Notion Business - Six months free of AgentMail We rely on these every day, and we tried to put together a package that helps you comprehensively for each part of the process of building and running software in AI. What comes with All Access - Everything in an existing paid Every membership: our daily writing, guides, camps, and software like Monologue, Cora, Sparkle, and Spiral - The Builder Pack, with more than $7,000 in partner offers - Unlimited email accounts use of Cora and unlimited Spiral usage - Members-only programming with me and the Every team and me Get All Access:

Dan Shipper 📧

183,025 views • 1 month ago

danshipper's profile picture

BREAKING: GPT-5.6 Sol is out—AND Codex has been merged into ChatGPT Desktop as ChatGPT Codex. This combo model and desktop app harness are the gold-standard for knowledge work in AI. 5.6 is powerful, fast, half the price of Fable, and my default for almost everything. We’ve been testing it internally Every 🪨 for about a month across coding, writing, design, and knowledge work. Here’s our day-zero vibe check: - An A-tier coder—but it’s not Fable. Sol scored 56/100 on our Senior Engineer benchmark compared to a 91 for Fable. I think the 56/100 undersells it, it's an excellent implementor, and very smart. But Fable just writes conceptually cleaner code and works better at the top end of task complexity. PRO-TIP: Use GPT-5.6 as Fable's subagent for the most goated combo in AI coding. - The best writer of the frontier models. It’s clearer and more concise than Fable or Opus 4.8, without the overexplaining or weird private language. It can one-shot marketing emails, help you workshop taglines, and explain complex concepts clearly. It's also super fast, which makes it easy to collaborate with. - Design is better, but not top-tier. It has noticeably more taste than 5.5, but Fable and Opus 4.8 are still playing at a different level. See examples in the video and vibe check below. - The real leap is knowledge work. Sol is the first model I’ve trusted to run whole loops of knowledge work—not just help with individual tasks. I use it to process email, surface decisions from meetings and Slack, find job candidates, scan Facebook Marketplace for furniture, and log my meals. It has shifted my job from doing the work to tending the system that does it. - The merged app is fine. I was extremely worried about this because I love the Codex app. OpenAI was caught in an interesting position: How to make an agent orchestration app for regular ChatGPT consumers, coders, and businesses all in one app. They now split the interface between ChatGPT Work and ChatGPT Codex. They're basically the same except Work hides code. And "Chat" has been demoted to 2nd tier status for quick questions in either one. It's not a big leap, but it's not a huge setback either. And it remains my favorite of the desktop agent orchestration apps. Verdict: If I really had to put my finger on it, I'd say Fable has way more big model smell. But that means it's a skill in itself to get value out of it—99% of people are still not there yet. GPT-5.6 is almost as powerful, but is easy to use, fast, and relatively cheap. It should give you an early sense of where model work is going. Full Every 🪨 Vibe Check:

Dan Shipper 📧

146,015 views • 1 month ago

danshipper's profile picture

BREAKING! Introducing Plus One: A hosted OpenClaw🦞 that lives in your Slack and comes pre-loaded with Every 📧's best tools, skills, and workflows. Set it up in one click, and use your ChatGPT subscription (or any other API key.) Bring your Plus One to work: Connected to the Every 📧 ecosystem Plus Ones automatically use Every 📧's agent-native apps, no setup required: - Cora for searching, sending, and managing email - Spiral for great writing in your voice - Proof ( for agent-native document editing Custom skills and workflows we use and love Plus Ones come pre-loaded with skills and workflows we use ourselves Every 📧 —some we've made, and some we think are great. - Content digest—summarizes the publications you read, starting with Every 📧 - Daily brief—your day's schedule and to-dos sent to you each morning - Animate—turn any static screenshot into an animation with Remotion - Frontend—Anthropic's front-end skill (which we use all the time!) We also make it fast to connect Google, Notion, Github, and more to your Plus One. Our goal is to give you a capable AI coworker right away, not a vanilla OpenClaw that you have to teach from scratch. Why we built Plus One OpenClaw🦞 has changed the way we work at Every. We effectively have a parallel org chart of AI coworkers, each with a name, a manager, and real responsibilities. Because of them our workflows are completely different—our company is different—and we would never go back. But getting here has been hard. Claws require a significant amount of manual setup and require a dedicated machine—like a Mac Mini—running 24/7 to stay responsive. We have learned that the hard part of Claws is the infrastructure around them—the hosting, the integrations, the skills, and the ongoing care. We’ve made them work great for our team, and we want to share everything we’ve learned with you. We're letting in 20 people a week to start, and scaling invites quickly from there. Every 📧 subscribers get priority. Bring your Plus One to work:

Dan Shipper 📧

260,028 views • 5 months ago

danshipper's profile picture

BREAKING NEWS: Anthropic just dropped Claude Ops 4.5!! It is by FAR the best coding model I've ever used. We've been testing it internally Every 📧 for the last few days, and it is an absolute paradigm shift for any kind of coding task. It extends the horizon of what you can vibe code The current generation of new models—Anthropic’s Sonnet 4.5, Google’s Gemini 3, or OpenAI’s Codex Max 5.1—can all competently build a minimum viable product in one shot, or fix a highly technical bug autonomously. But eventually, if you kept pushing them to vibe code more, they’d start to trip over their own feet: The code would be convoluted and contradictory, and you’d get stuck in endless bugs. We have not found that limit yet with Opus 4.5—it seems to be able to vibe code forever. Takes working in parallel to a whole new level because it's far better at planning and coding, it can work with more autonomy—meaning you can do more in parallel without breaking anything . Kieran Klaassen worked on 11 different projects in six hours—and had good results on all of them. Great at design iteration Opus 4.5 is incredibly skilled at iterating through a design autonomously using an MCP like Playwright. previous models would lose the thread after a few cycles, or say a design was done when it wasn't. Opus 4.5 is incredible at autonomously iterating until a design is pixel perfect. we have a full 4,000 word vibe check on Every 📧 right now with everything we tested:

Dan Shipper 📧

272,699 views • 9 months ago

danshipper's profile picture

BREAKING: GPT-5.5 "Spud" is out and it is a BEAST We've been testing it Every 📧 for the last 3 weeks on everything from coding, to writing, to knowledge work. Here's our day 0 vibe check: - It's a step change in coding AND it's easy to talk to. It's fast and friendly and quickly became my daily driver. But it's also a coding powerhouse—a really rare combination. - It scored 62/100 on our Senior Engineer benchmark. Opus 4.7 scored only a 33/100. (But GPT-5.5 performed best when using an Opus 4.7 plan). Naveen Naidu used over 900 million tokens during testing—and it let him ship production features for Monologue at both high speed and quality. - It has serious conceptual clarity. It can hold a complex plan in its head over hours of work, without getting distracted by existing code. This makes it the first model that we've tested that can perform well on complex refactors requiring deleting and reimagining an substantial existing codebase. - It's a very good writer. This is the first OpenAI model in about a year that got our writers Every 📧 to switch away from Claude. 5.5 has Katie Parrott's seal of approval—not an easy task. Its writing feels more organic and it's better at mimicking a writing style without going overboard. - It's great for agentic knowledge-work. This is the first OpenAI model that manages to be both a stellar senior engineer AND that can be used for everything from spreadsheets to research. It's crazy fast, and it's amazing inside of the Codex desktop app, and got much of our team to switch away from Claude Code and Cowork during the testing period. However, it's not a perfect model. - 5.5 still loses to Opus 4.7 on plan quality. It's plans are extremely readable but Opus has better attention to detail and sharper insight. - 5.5 still loses to Opus 4.7 by a bit on front-end and full-stack product work. Kieran Klaassen found that it wasn't quite as good when full-stack thinking and design are involved. And it's not great writing Ruby. - 5.5 is a great vibe coder but if you're vibe coding without a plan it's worse than Opus. Mike Taylor found that Opus is better at reading in between the lines on underspecified vibe-coding tasks. Overall GPT-5.5 is a massive achievement from OpenAI and it deserves a serious look as your daily driver. Read our full vibe check on Every 📧 here:

Dan Shipper 📧

130,382 views • 4 months ago

danshipper's profile picture

Andrew Wilkinson (Andrew Wilkinson) has been waking up at 4 a.m. because he can’t stop building with Anthropic’s Opus 4.5. He started vibe coding a couple of years ago, but it felt like the Palm Treo era of the smartphone—exciting, but not quite there. You could generate an app, but it would get stuck in bug loops or break the moment you pushed it further. Then he tried Opus 4.5 in Claude Code. It felt, he says, like having a “$100,000-a-month payroll of engineers” working for him 24/7. He’s built practical AI automations into every corner of his work and life, including: - A relationship counselor app called Deep Personality that consolidates 20 clinically validated personality tests into a 40-minute assessment, then generates a 45-page analysis. When both partners complete it, it maps compatibility and predicts conflicts—Wilkinson says it laid out every fight he and his girlfriend have. - A custom email client he built by handing Claude Code his Gmail credentials and describing his ideal workflow. It triages emails by priority and sender, handles quick replies via multiple choice, and walks him through complex emails question by question before drafting. - A personal stylist that texts him four outfit recommendations every morning. It checks the weather, pulls from a spreadsheet of his entire wardrobe (photos converted to CSV by Claude), generates four outfit options rendered as images with Nano Banana 2, and texts him what to wear down to the watch. - A Lindy agent that acts as an AI referee of sorts—it records his meetings and texts him if it detects psychological red flags like manipulation or gaslighting. The bar is high—he only gets a notification every few months—but when he does, it usually confirms a gut feeling he already had. Andrew is the cofounder of Tiny, the holding company that owns businesses like AeroPress and Dribbble. Earlier in his career, Andrew was a web designer, and he fits one of my predictions for 2026: Designers, who know how to create great experiences for users, are the unsung group most empowered by this AI moment. I had him on Every 📧's AI & I to talk about Opus 4.5, what he’s building with it, and how it’s changing the way he thinks about acquiring software businesses at Tiny. This is a must-watch for anyone who wants to put AI to work in their day-to-day life. Watch below! Timestamps: Introduction: 00:01:07 Why Opus 4.5 feels like the iPhone moment for vibe coding: 00:02:48 Why designers have a unique advantage with AI: 00:08:31 How Andrew built a custom email client with Claude Code: 00:14:10 An AI trained on your relationship that predicts your fights: 00:18:13 Using AI meeting notes to make your life better: 00:30:40 Don't inject your opinion into prompts: 00:35:11 Andrew's Claude Code tips and workflows: 00:40:21 Your personal stylist is a prompt away: 00:47:59 How AI is changing the way Andrew invests in software: 00:53:17

Dan Shipper 📧

154,567 views • 7 months ago

danshipper's profile picture

In the future, you’ll be able to accomplish a goal by just giving Claude an outcome and a budget. That’s the direction Anthropic is building in with its new Managed Agents features, announced at this week’s Code with Claude developer event. The basic idea: Claude, wrapped in a computer in the cloud, that you can spin up, scale, and manage as needed. Anthropic is taking on the infrastructure that kills most agent products, and making sure that it scales to meet the needs of agents running 24/7. On this week’s AI & I from Every 📧, I talk with Angela Jiang (Angela Jiang), head of product for the Claude platform, and Katelyn Lesse (Katelyn Lesse), head of engineering for the Claude platform, about what Anthropic is building and what it takes to make agents reliable in production. We get into: - Why the "build a generic harness, hot-swap any model behind it" playbook is already outdated. Angela points to eval data on Memory where the same task across different harnesses performed drastically differently. - The infrastructure wall every team hits in production—and why Katelyn thinks “my sandbox died and took the agent with it” is the real reason internal agents don't ship. - Why Anthropic is so bullish on using file systems and skills within Claude, including Angela's argument that those early design choices can compound for years. This is a must-watch for anyone trying to take an agent past the demo and into production. Watch below! Timestamps: How the Claude platform evolved from API to agents: 00:01:48 The primitives that make up Claude Managed Agents: 00:04:09 Why the harness and the model are becoming a single unit: 00:10:37 The infrastructure wall that kills most agent projects in production: 00:18:49 Why team agents need a different shape than individual productivity tools: 00:24:49 How Anthropic's legal team uses an agent to review marketing copy: 00:26:36 Using multi-agent orchestration for advisor strategies, adversarial pairs, and swarms: 00:34:24 How to measure agent success with outcome and budget as the end state: 00:35:50 What the platform looks like a year from now, when Claude writes its own harness: 00:39:11

Dan Shipper 📧

66,339 views • 3 months ago