Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

IT'S NOT THAT HIS FABLE 5 IS MORE CAPABLE THAN YOURS. IT'S THAT HE LETS IT KEEP WORKING. MOST PEOPLE BUY AN AI BUILT TO OPERATE FOR DAYS THEN SHUT IT DOWN AFTER A FEW MINUTES. Nine out of ten users open a chat, assign one task, and close...

20,116 Aufrufe • vor 16 Tagen •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

ANTHROPIC'S PRODUCT CHIEF HAS USED CLAUDE FABLE 5 FOR MONTHS BEFORE ANYONE ELSE. HERE'S WHAT HE LEARNED ABOUT THE MOST POWERFUL MODEL YET Mike Krieger co-founded Instagram and now runs product at Anthropic. He's had Claude Fable 5 for two months before the public, and his takeaway is that it changes how you have to work, not just how much you get done. Here's what stood out, and what to actually do with it 1. It holds the whole project, so stop chopping tasks small. The old habit was breaking work into model-sized pieces and stitching them. Fable keeps the whole thing in context. What to do: stop pre-slicing your prompts into tiny steps. Hand it the full goal and the intent behind it, the way you'd brief a senior engineer, and let it sequence the work itself 2. Delegate big, async, and overnight. He sets it on a hard task at night and wakes to it finished, including the model getting itself unstuck when a service died, scaffolding a workaround, and documenting it. What to do: stop babysitting one prompt at a time. Kick off long jobs and walk away. Run several sessions at once instead of one you watch 3. The skill is planning now, not typing. His day moved to long architecture conversations up front, then execution in chunks. What to do: spend your first prompts planning, not building. Then ask it to output an HTML page or markdown doc of the plan so your team aligns before any code is written. That early alignment is the new leverage 4. Match the effort level to the task. Fable's range is wide, so a heavy reasoning pass on a tiny UI tweak is overkill (and pricey). What to do: dial effort down for small jobs, save the deep thinking for hard ones. And don't use your most expensive model for quick questions, keep a fast model for those 5. Verification is the real bottleneck now. The hard part isn't getting output, it's trusting it. What to do: make every change ship with proof. Have Claude attach a screenshot or video of what it built, so you can see the result instead of reading the diff. Then stand behind the decisions yourself before you merge 6. Cost is per-result, not per-turn. Fable is expensive per call but often one-shots what other models need ten turns to get right. What to do: judge cost by what it takes to finish the task to your satisfaction, not the price of a single message. Give it a real task and see how far it gets before you jump in His bigger point: software engineering isn't over, it's different. The craft moved from writing code to owning intent, taste, and what actually ships. The floor rose so anyone can build, and the ceiling rose so experts go further than before Bookmark this

Yarchi

30,743 Aufrufe • vor 1 Monat

One guy keeps a farm of Mac minis on his desk and says each $600 box brings him $2,000 a month while he sleeps. AND THE HARDWARE ACTUALLY WORKS. But the number is not even the interesting part. The broken part is HOW: his AI no longer sits in a chat window. It sees the screen, moves the mouse itself, types and clicks the interface like a human at a computer. That is it. While most people still run AI in a chat and ask it for text, he sat Claude down right at the computer and put it to work with its hands. He automated not a single task but the workplace itself. How it actually works: on every Mac mini Claude runs with computer use turned on and the official Claude API docs spell it out: screenshot capture, mouse control, keyboard input, desktop automation. The agent opens the browser and the apps itself and runs the boring routine on a schedule: pulls leads, fills the CRM, checks orders, runs QA on the site. One box, one quiet worker that does not sleep and does not ask for a salary. His math is simple: a Mac mini is $600 once, Claude Max is $200 a month, and a live white-collar worker on the same routine costs a business $4,000 and up. So he rents out each node to a client as an AI worker for about $2,000 and 6 Mac minis come out to around $12,000 a month with costs a bit over $1,000 on subscriptions. But the $12,000 is his projection not a revenue dashboard: the video has no client, no task log, no working automation at all. The real asset here is not the stack of hardware but the one repeatable process the agent actually closes. Because a Mac mini on its own earns nothing. The money shows up exactly where the boring browser routine used to be done by hand for a salary and now you can hand it to an agent for the price of a subscription. Computer use is still in beta, almost nobody builds a service on it and the demand for cheap GUI routine is huge. The window is open for literally the next few months. Most people will watch this, laugh at the "$600 AI worker" and close it. And the ones who actually put an agent on one boring task and grind it into a repeat will ride this wave while it is still empty. Would you sit an AI right at your own computer on the boring routine or are you still clicking through it by hand?

Sorven

11,948 Aufrufe • vor 1 Monat

To be a successful founder, you have to believe that what you're working on is going to work — despite knowing it probably won't! That sounds like an oxymoron, but it's really not. Believing that what you're building is going to work is an essential component of coming to work with the energy, fortitude, and determination it's going to require to even have a shot. Knowing it probably won't is accepting the odds of that shot. It's simply the reality that most things in business don't work out. At least not in the long run. Most businesses fail. If not right away, then eventually. Yet the world economy is full of entrepreneurs who try anyway. Not because they don't know the odds, but because they've chosen to believe they're special. The best way to balance these opposing points — the conviction that you'll make it work, the knowledge that it probably won't — is to do all your work in a manner that'll make you proud either way. If it doesn't work, you still made something you wouldn't be ashamed to put your name on. And if it does work, you'll beam with pride from making it on the basis of something solid. The deep regret from trying and failing only truly hits when you look in the mirror and see Dostoevsky staring back at you with this punch to the gut: "Your worst sin is that you have destroyed and betrayed yourself for nothing." Oof. Believe it's going to work. Build it in a way that makes you proud to sign it. Base your worth on a human on something greater than a business outcome.

DHH

96,510 Aufrufe • vor 1 Jahr

I’ve been using GPT-5.6 Sol internally for the past two months, I've spent probably 25+ billion tokens. Here’s my review and comparison to Fable 5: > Let's start with the analogy because everyone seems to be giving theirs - GPT-5.6 is likely the last version of the GPT-5 training run series. It's kind of like an athlete at their peak. Through years of experience in the game, they've become the most reliable player and has the highest game IQ. But, there's no more room to grow. Fable on the other hand, being essentially the first version of a new training run, is the first round draft pick rookie. Raw talent mixed with the energy only a young person would have results in some incredible plays we didn't think possible, but also mistakes due to lack of experience. But that rookie will only improve and likely will be better than the veteran ever was because it's a new game and a new era. > GPT-5.6 is genuinely better at long, sustained work. With /goal, I've had it running complex projects for days with almost no intervention. It built a Minecraft-style game, kept adding features and mobs after the core game worked, and only stopped because I stopped the run. I never felt as though I had to jump in and guide it back to the right path. > It keeps finding useful work when you give it a concrete finish line. I had it recreate Excel with a loop. It inspected the real desktop excel app with Computer Use, comparing that against its own build, and closing the gaps. I stopped it after six days after it had built an incredible amount of functionality. > It's faster than other models in two different ways. The raw generation speed is higher, something OpenAI has been putting effort into. But it also takes a shorter path to solutions. It wanders less, changes less code, and generally knows how to get things done directly. In daily use, it feels about 2-3x times faster than Fable. That's my impression, not a controlled benchmark. The difference is large enough that I notice it constantly. > It works well across a wide range of tasks. I use it for one-line edits, quick questions, browser chores, and multi-day builds without changing my prompting style. Speaking of browser control, its the best ever I've used. To the point where I actually use it often. If a task lives on a website, GPT-5.6 usually opens the browser and does it there instead of asking for an API key or forcing everything through the terminal. When I switched back to GPT-5.5, it went straight to the command line even when the browser was clearly the better tool. > And it can handle real browser work, not just toy demos. During a data import, I had it monitor Supabase and resize instances as the load changed. It stayed on the dashboard, adjusted capacity, and checked the result without an API or a custom script. > I also gave it a full Google Workspace migration. It moved Forward Future from to preserved the old aliases, and configured MX, SPF, and DKIM. Before a consequential save, it stopped, explained exactly what would change, and waited for confirmation. > The reasoning setting matters a lot. Light is good for questions and small edits. High and Extra High are the sweet spots for serious work. Ultra usually takes longer than the extra thinking is worth and burns tokens. > I love that 5.6 is split into 3 sizes. Not only can you control speed and cost that way, but you still also have the thinking effort setting for each of them. Very precise controls. I just wish Codex automatically routed my prompts for me. > Its personality is blunt and a little bland. Claude feels warmer and more natural to talk to. GPT-5.6 is more clinical, but I like that for work. It gives me enough explanation and rarely pads the answer. I usually have to ask Fable to explain things more simply and/or more concise. > Its front-end taste has improved, but the default is predictable. Left alone, it turns websites into PowerPoint decks with huge statements and hard section breaks. The good news is that it takes design direction well and can revise without destroying the parts that already work. > It still makes confident mistakes. I asked it to rebuild parts of a system, and it told me the job was finished. Later, I found out it wasn't. Bits of its internal process also leak into the answer occasionally. > Claude Fable is more naturally autonomous on large, open-ended projects. GPT-5.6 is easier to reach for. I don't need to invent a huge project to justify using it. It works just as well for a small edit or browser chore. > GPT-5.6 is also cheaper. Sol costs $5 per million input tokens and $30 per million output tokens. Fable costs $10 and $50. Cached input is cheaper too. Still, cost per finished task matters more than cost per token. > GPT-5.6 isn't the best at everything, and it still needs supervision. But it generates faster, wanders less, works at almost any scale, and wastes less of my time. It's the model I have the most confidence in to get the job done right the first time. I put together a full breakdown with all the tests, prompts, and examples on a site. You can read it here:

Matthew Berman

186,712 Aufrufe • vor 24 Tagen

BREAKING: GPT-5.6 Sol is out—AND Codex has been merged into ChatGPT Desktop as ChatGPT Codex. This combo model and desktop app harness are the gold-standard for knowledge work in AI. 5.6 is powerful, fast, half the price of Fable, and my default for almost everything. We’ve been testing it internally Every 📧 for about a month across coding, writing, design, and knowledge work. Here’s our day-zero vibe check: - An A-tier coder—but it’s not Fable. Sol scored 56/100 on our Senior Engineer benchmark compared to a 91 for Fable. I think the 56/100 undersells it, it's an excellent implementor, and very smart. But Fable just writes conceptually cleaner code and works better at the top end of task complexity. PRO-TIP: Use GPT-5.6 as Fable's subagent for the most goated combo in AI coding. - The best writer of the frontier models. It’s clearer and more concise than Fable or Opus 4.8, without the overexplaining or weird private language. It can one-shot marketing emails, help you workshop taglines, and explain complex concepts clearly. It's also super fast, which makes it easy to collaborate with. - Design is better, but not top-tier. It has noticeably more taste than 5.5, but Fable and Opus 4.8 are still playing at a different level. See examples in the video and vibe check below. - The real leap is knowledge work. Sol is the first model I’ve trusted to run whole loops of knowledge work—not just help with individual tasks. I use it to process email, surface decisions from meetings and Slack, find job candidates, scan Facebook Marketplace for furniture, and log my meals. It has shifted my job from doing the work to tending the system that does it. - The merged app is fine. I was extremely worried about this because I love the Codex app. OpenAI was caught in an interesting position: How to make an agent orchestration app for regular ChatGPT consumers, coders, and businesses all in one app. They now split the interface between ChatGPT Work and ChatGPT Codex. They're basically the same except Work hides code. And "Chat" has been demoted to 2nd tier status for quick questions in either one. It's not a big leap, but it's not a huge setback either. And it remains my favorite of the desktop agent orchestration apps. Verdict: If I really had to put my finger on it, I'd say Fable has way more big model smell. But that means it's a skill in itself to get value out of it—99% of people are still not there yet. GPT-5.6 is almost as powerful, but is easy to use, fast, and relatively cheap. It should give you an early sense of where model work is going. Full Every 📧 Vibe Check:

Dan Shipper 📧

145,238 Aufrufe • vor 24 Tagen

i watched gemma 4 12b build something genuinely impressive today, and then loop itself to death right in front of me. the full run is in the video, sped up but completely uncut, watch it to the end and you will catch the exact moment it stops building and starts looping right in the middle of the work. the task was clean, build a single file gravity simulator, n-body physics, orbits, collisions, running locally on one 3090 through an agent. and for ten minutes it was a joy to watch. it reached for a symplectic integrator on its own, the correct one, the kind that keeps orbits stable instead of spiralling out. real gravity with softening, proper orbital velocities, momentum conserved on collision. the physics was right. the thing actually worked. then on the very last step, writing a few tests to prove its own code, it fell into a loop. not a crash, a loop. it started repeating itself and would not stop. ten more minutes, thirty four thousand tokens into a single answer, the same fragments over and over, until i killed it myself. so it's not that gemma can't code. it did the hard part beautifully. it cannot finish. it cannot hold a long task together without unravelling, and finishing is the entire job in agentic work. here's the part that stings. i run this exact task, same harness, same card, on the chinese open models, qwen especially, and i never see this. they build it, they test it, they stop. every single time. google has the raw capability, you can see it sitting right there in the code, and then the model loops itself to death on a task a 27b from alibaba finishes clean. open weights, apache 2.0, so much to love on paper. i just need it to know when to stop talking.

Sudo su

39,574 Aufrufe • vor 1 Monat