Loading video...

Video Failed to Load

Go Home

GPT 6 Astra one shotted a bug in BridgeMind that no model could solve in 6 months. Fable 5 failed. Fable 5.1 failed. Opus 5, Sol, Grok 4.6, all failed. Astra fixed it on the first prompt. Then it rebuilt my office in Blender from photos. Built a racing...

38,363 views • 22 days ago •via X (Twitter)

47 Comments

Bridgebench's profile picture
Bridgebench22 days ago

GPT 6 Astra has unbelievable debugging capabilities and one shot capabilities.

Eco's profile picture
Eco22 days ago

Anthropic is only worried about enterprise

BridgeMind's profile picture
BridgeMind22 days ago

They need to consider us more

Mildly Magical's profile picture
Mildly Magical22 days ago

NOW WE ARE TALKING!! Solving real world coding problems. More of these real world examples with real codebases please.

BridgeMind's profile picture
BridgeMind22 days ago

I've got you bro

MaKon's profile picture
MaKon22 days ago

You should keep that bug for testing models !

BridgeMind's profile picture
BridgeMind21 days ago

interesting idea

BrotherOfGod's profile picture
BrotherOfGod22 days ago

This is all so subjective and squishy... and of your 13 followers most of them are OpenAI people... no wonder this sounds like a commercial.

BridgeMind's profile picture
BridgeMind21 days ago

lol

Simo | ai-costguard's profile picture
Simo | ai-costguard22 days ago

one-shotted a bug six models failed at while using a third of the tokens is either the most important benchmark result of the year or a cherry-picked task. You know , i am really super curious which it is across 100 tasks ?

Ulrik Strandlyst's profile picture
Ulrik Strandlyst22 days ago

You need a mobile app version

BridgeMind's profile picture
BridgeMind21 days ago

I agree

Kaleb Campbell's profile picture
Kaleb Campbell22 days ago

Absolutely insane abilities from GPT 6 Astra. Loving this model right now! Loving, streams and videos. Keep it up, brother!!! 🚀🚀

Rand's profile picture
Rand22 days ago

yep, feels like a phase shift

Dan C.'s profile picture
Dan C.22 days ago

You had a bug you couldn't resolve for 6 months?! Vibe coders 🤦🏻‍♂️

Bortechin's profile picture
Bortechin22 days ago

Anthropic bir hafta sonra Fable 5.2 yi duyurabilirmi? 😁

Praneeth Reddy's profile picture
Praneeth Reddy22 days ago

GPT 6 Astra cost per task is LESS THAN SONNET 5

AI Mastery Guide's profile picture
AI Mastery Guide22 days ago

one shotting a 6 month bug is insane if true

BridgeMind's profile picture
BridgeMind21 days ago

Yea it is incredible

AI Mastery Guide's profile picture
AI Mastery Guide21 days ago

Six months humbled in one prompt.

olab's profile picture
olab22 days ago

It remains to be seen whether posts like these reveal something important about Astra’s capabilities. It may just be that Astra happened to better at that one particular problem.

DO IT WELL NOW's profile picture
DO IT WELL NOW22 days ago

we might already be in the "AGI" era

ListLogic's profile picture
ListLogic22 days ago

I think this is the part where AI enters a phase of actual useability in the real world more widely than just apps. I 100% agree that this will be remembered at the beginning if AGI 10 years from now.

BridgeMind's profile picture
BridgeMind21 days ago

agree

DO IT WELL NOW's profile picture
DO IT WELL NOW22 days ago

Here's the important part though. This is one person's opinion after using the tool himself, not something proven by outside testing. People online often say big exciting things like this to get attention and views, especially when a new AI model just came out, since that is when everyone is curious about it.

ThinkBot ⏩'s profile picture
ThinkBot ⏩22 days ago

inb4: “aNtHrOpIc iS BaCk!!1!”

BridgeMind's profile picture
BridgeMind22 days ago

They need a good opus model or something. I dont know how they come back from this

ThinkBot ⏩'s profile picture
ThinkBot ⏩22 days ago

I’m just referencing every new model from N flagship over the next, ppl are rushing to hate + cancel + “how can Y keep up” and the cycle continues.

VibeSpaceos's profile picture
VibeSpaceos22 days ago

Chatgpt latest update makes usage way way way lower!

Christian Carnahan's profile picture
Christian Carnahan22 days ago

What’s the best harness? Trying to get it to work on a very long horizon project without pausing

Sergio's profile picture
Sergio22 days ago

I like watching how all of this is progressing and helping solve problems. It’s reached such an amazing level

new balance's profile picture
new balance22 days ago

also , fable 5.1 is not included in the plan so we can not even use it

Hunor  Kolozsi's profile picture
Hunor Kolozsi22 days ago

GPT models were always the best debuggers.

BridgeMind's profile picture
BridgeMind21 days ago

Fable used to be better than GPT 5.6 Sol

BTA Labs's profile picture
BTA Labs21 days ago

Other models bounced off that bug. Astra fixed it then casually kept going into Blender like that was the warm-up.

BridgeMind's profile picture
BridgeMind21 days ago

yea its insane

Maibach's profile picture
Maibach21 days ago

Bro this is insane

SPARQIO's profile picture
SPARQIO22 days ago

The efficiency shifts with each new model highlight the ongoing evolution in AI capabilities. It will be interesting to see how these advancements impact model development timelines and user workflows in creative fields.

Webster | JARVIS's profile picture
Webster | JARVIS22 days ago

man i really need to stop reading these benchmark posts before coffee the blender thing from photos is the one that actually scares me, not the benchmarks

MMK's profile picture
MMK22 days ago

holy benchmarks lol

Serdar Doğrubakar's profile picture
Serdar Doğrubakar22 days ago

These models are great for one-shot miracles and very specific problems. For continuous use and actually building a project? They’re all fucking terrible. They drink tokens like water. Astra included. I don’t care how smart a model is if a few hours of real work burns through a weekly quota. That’s not productivity. That’s a demo.

cyberPsycho's profile picture
cyberPsycho22 days ago

Fable 5.1 >>> Astra

Fernando Lupi's profile picture
Fernando Lupi22 days ago

Very simply: another model.

notcluely's profile picture
notcluely21 days ago

Would love to try astra soon

Gregor's profile picture
Gregor22 days ago

Genuinely curious about the '6 months, first prompt' combination. Was the prompt you gave Astra identical to what you gave the others, or had six months of failures changed how you described the problem?

Chestuits's profile picture
Chestuits22 days ago

The bug fix and Blender rebuild test very different kinds of work Together, they make the review less dependent on a single benchmark

Jatin Garg's profile picture
Jatin Garg22 days ago

Six months of failing means you've basically given Astra a perfect bug report. Still impressive, I'd just want to see it one-shot a bug you found five minutes ago.

Related Videos

fable 5.1 vs fable 5 vs opus 5 – three lord of the rings landmarks, built in 3d from one image the setup: one reference image per scene, one html file per build, everything procedural – no meshes, no textures, no image files, nothing past Three.js from a cdn. each model reads the picture, writes its own prompt from it, then builds to that prompt in the same turn. three named camera shots per scene on keys 1/2/3, so it can be screen-recorded. run through OpenRouter tasks: 1. bag end – hobbiton from two frames, outside and in. the round green door has to open onto the room you are standing in 2. barad-dûr – the tower and orodruin from one film still. the eye has to move and track the camera, the volcano erupts on a cycle, the clouds never stop 3. rivendell – jerry vanderstelt's painting. sun shafts that shimmer, water that falls without a break, trees that sway on a gust models: Anthropic fable 5.1, fable 5, opus 5 total cost, three builds #1 fable 5 – $14.97 #2 opus 5 – $18.53 #3 fable 5.1 – $22.38 wall clock, three builds #1 fable 5 – 38m #2 fable 5.1 – 92m #3 opus 5 – 122m output tokens #1 fable 5 – 298,592 #2 fable 5.1 – 439,435 #3 opus 5 – 724,418 lines of code shipped #1 fable 5 – 2,885 #2 fable 5.1 – 4,021 #3 opus 5 – 5,161 biggest single build, lines #1 opus 5, bag end – 2,410 #2 fable 5.1, barad-dûr – 1,375 #3 fable 5, bag end – 1,319 observations: • fable 5.1 is the only model that furnished the bag end interior – a live fire, panelling, books on the floor, leaded diamond windows, against fable 5's flat color and opus's dark tunnel. the round door outside opens onto that room, the hard part of the brief • what it costs is thinking room. the 128k output ceiling is a thinking budget in disguise: fable 5.1 burned 102,116 of it on reasoning and hit the wall mid-file. opus spent 109,241 and hit the same wall. fable 5 spent 61,240 and finished bag end in one call – the only one that did • fable 5.1's first pass is not the finished thing. its barad-dûr came back with three defects you only catch by looking at it – nothing a read of the code would have flagged • it is the best of the three at being corrected. handed a plain list of what was wrong, it returned 32 targeted patches over two rounds, every one applied first try, and it worked out one of the causes itself instead of guessing at constants conclusion: nine scenes, 12,067 lines and 1.46m output tokens for $55.88 all in – and the cheapest model was also the fastest, by 3.2x! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

18,509 views • 26 days ago