正在加载视频...

视频加载失败

GPT 6 Astra vs SWE 2.0 tested both models with same prompt at highest reasoning available > astra took 80 minutes to complete this task and costed $64 > swe 2.0 took 25 minutes and costed zero dollars really surprised by swe results here and astra looks nerfed right...

30,824 次观看 • 1 天前 •via X (Twitter)

51 条评论

J A Z I I 的头像
J A Z I I1 天前

you can play around here: both models were test in devin desktop app

Damianonyx ☯︎ 的头像
Damianonyx ☯︎1 天前

I'd chosen SWE at first glance, but GPT-6 Astra did it better. Compare GPT-6 Astra to Nex-N2.5-Pro. It's currently free on OpenRouter.

J A Z I I 的头像
J A Z I I1 天前

lemme run same prompt on next will do tonight

Damianonyx ☯︎ 的头像
Damianonyx ☯︎1 天前

Okay, please tag me when you do 🙏🏽

martwypoeta 的头像
martwypoeta1 天前

asrra 😭

J A Z I I 的头像
J A Z I I1 天前

it shows real human edited lmaoo

Shubh 的头像
Shubh1 天前

swe 2.0 did an incredible job

Parody Elon Musk (Crypto Edition) 的头像
Parody Elon Musk (Crypto Edition)1 天前

astra is a scam

AshutoshShrivastava 的头像
AshutoshShrivastava1 天前

How zero? Is it free to use?

J A Z I I 的头像
J A Z I I1 天前

yes via Devin

am.will 的头像
am.will1 天前

swe cooked here

samyok 的头像
samyok1 天前

this is insane

J A Z I I 的头像
J A Z I I1 天前

you guys cooked, cognition w

Tim Jayas 的头像
Tim Jayas1 天前

is swe 2.0 Kimi-k3.1 ?

J A Z I I 的头像
J A Z I I1 天前

you could say that but it's free

GBE 的头像
GBE1 天前

its soo good tho

levithefirst 的头像
levithefirst1 天前

I thought Astra was end game tf is going on

J A Z I I 的头像
J A Z I I1 天前

Astra is nerfed pretty hard rn

levithefirst 的头像
levithefirst1 天前

lol

painn 的头像
painn1 天前

wait $0 for this?😭 swe 2.0 won

J A Z I I 的头像
J A Z I I1 天前

free in devin for now

painn 的头像
painn1 天前

without a devin plan? (i was gifted a max btw)

J A Z I I 的头像
J A Z I I1 天前

yeah try on that without plan it's not available anywhere except Devin for now

WolfSWAP | SWAP & WIN 的头像
WolfSWAP | SWAP & WIN1 天前

how did free model cooked astra?

J A Z I I 的头像
J A Z I I1 天前

I wanna know too

Miguel 的头像
Miguel1 天前

Let's hope they can keep up the capacity and not having to nerf like others do

John Kler 的头像
John Kler1 天前

Oh wow. Swe 2.0 is absolutely incredible. Didn't expect this from a relatively small product/lab.

J A Z I I 的头像
J A Z I I1 天前

cognition is coming for top models

m0h 的头像
m0h1 天前

which model is SWE 2.0???

J A Z I I 的头像
J A Z I I1 天前

swe 2.0 is model name made by cognition available in Devin for free via their subs

m0h 的头像
m0h1 天前

oh not a fan of Devin, where does their sub start from

J A Z I I 的头像
J A Z I I1 天前

20 bucks

Shikhar 的头像
Shikhar1 天前

is swe 2.0 free in devin?

J A Z I I 的头像
J A Z I I1 天前

yeah free in devin that's why zero

Berna 的头像
Berna1 天前

Bro AI is becoming scary advanced

K2S 的头像
K2S1 天前

Damn u used 64$ for this output and swe 2.0 😱

J A Z I I 的头像
J A Z I I1 天前

yeah man, I have to test kek

K2S 的头像
K2S1 天前

Sure test but this type of test 🤣 🤣

Inferno 的头像
Inferno1 天前

This is insane bro

Veee 的头像
Veee1 天前

the way you put out output is diabolical dude

J A Z I I 的头像
J A Z I I1 天前

haha 😂😆

kepo 的头像
kepo1 天前

SWE 2.0 underrated

J A Z I I 的头像
J A Z I I1 天前

you should try it, it's free

am.will 的头像
am.will1 天前

costed isn't a word btw (at least not in this context) :P

J A Z I I 的头像
J A Z I I1 天前

if you can understand it's good, doesn't matter the wording, shows ai didn't write it kek

Qwerty 🧀 的头像
Qwerty 🧀1 天前

Astra don't have competition now

Gregor 的头像
Gregor1 天前

The $0 makes it hard to call this a fair comparison. Speed and cost gaps at highest reasoning usually mean different step counts, not a broken model. Output diff would tell you more than the timer.

Cedric 的头像
Cedric1 天前

There's no actual API conversion for SWE 2.0 hence $0? I get that it's free but there's no actual conversion if it were to be in api cost?

J A Z I I 的头像
J A Z I I1 天前

yeah i think so

Fayeque 的头像
Fayeque1 天前

How is the performance(speed and quality) compare to deepseek v4.1 flash?

Gunbit 的头像
Gunbit1 天前

80 minutes is brutal. Astra's reasoning depth is impressive but the token economics don't scale for production workflows. Have you tested it with smaller, decomposed prompts? Often cuts time by 60% without quality loss.

相关视频

I GAVE GPT-6 ASTRA AND GROK BOT $1,000 EACH AND THE BETTER TRADER F*CKING LOST grok closed $1,742.15. astra closed $1,634.27. eleven points apart same day, same page, six agents each. one order: beat the other room or i cut funding astra was the better trader on every line people brag about → coin side: astra $586.42, grok $468.90 → astra went 17 of 24 green, grok went 16 of 24 → astra caught the MARLIN low at 57.4K and the 818K top, +$462.44 on two tickets → astra took 92% of its day out of that one coin → grok took 37% of his day out of prediction books, +$273.25 → astra ran the same books for $47.85 so how did the worse trader take the money? a calendar apple held a keynote that day. grok bought IPOD at 559K into it, +$306.48, biggest ticket of the duel. then he traded the market pricing it wrong: "apple foldable iphone by sep 30", NO at 45c, closed at 97.9c both opened the same bitcoin book over 80k. astra took YES at 34.5c and watched it bleed to 13.5c, its worst book of the day. grok took NO at 58.5c and closed it at 93.5c same market, same hour, opposite sides. no benchmark tells you that. leaderboards score answers, not tickets both floors trade on one page anyone can open: a manhattan seat that reads a calendar costs $180,000 a year. mine cost less than dinner and argued all night the six-role setup i handed both sits in the article i quoted, free you can rent the smarter model, not the habit of reading tomorrow whole duel is in the clip below, tick by tick. bookmark it before the feed eats it which one would you fund after that scorecard?

savip.

34,908 次观看 • 1 天前

gpt astra vs fable 5.1 at goldberg machine gpt 6 astra – openai, landed on OpenRouter less then hour ago, provider pinned to openai fable 5.1 – anthropic, shipped sep 1 we put the two models on one job: a rube goldberg machine in three.js that presses a button and detonates a bomb the setup: one self-contained html file, three.js from a cdn, everything else procedural – no textures, no models, no physics engine, every collision hand-written. the hard part sits in the brief: a domino may only fall once the previous one actually touches it, checked by real overlap every frame, never by a timer. same rule for the hammer hitting the button and the button firing the bomb. one continuous camera, its speed driven by whatever is moving. we recorded both scenes frame by frame – 1200 frames, 60 fps, exactly 20 seconds – and stepped both by hand to read the telemetry. - cost #1 astra – $1.84 #2 fable – $29.16 - time #1 astra – 9m 56s #2 fable – 1h 12m - tokens #1 astra – 45k #2 fable – 360k - lines of code astra – 881 fable – 744 observations: • we told it what we saw and nothing else – no diagnosis, no patch. we never edit a model's code. round two ran the whole chain to the blast. • both files are deterministic. two runs each, identical state to twelve decimals, and neither model reached for math.random. conclusion: 15.8x cheaper and 7.2x faster, and it still took a second round to get the ball into the bucket! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

37,131 次观看 • 7 天前

36 GROK AGENTS. ASTRA ON FREE CREDITS. DUAL-MODEL ROUTING IS THE EDGE. not one chat window burning a paid invoice all day a swarm that routes cheap work to grok and only wakes gpt-6 astra when the task is actually hard ▹ the stack 36 agents on grok bot with flexible settings per role monitor, plan, write, code, QA, ship, each with its own lane part of the fleet is wired to gpt-6 astra through a china free-credit service layer free credits are not a toy promo here they are the fuel for frontier spikes without a monthly bleed ▹ dual-model routing easy jobs stay on grok: speed, volume, always-on loops hard jobs jump to astra: reasoning, long builds, sharp code the router decides by task type, not by ego if astra is not needed, astra does not spend free-credit bursts buy the expensive brain grok agents keep the factory running between bursts that is how the system feels unlimited not by breaking quotas, by refusing to waste them ▹ why it hits different most people pay frontier prices for every mid task operators split the brain and protect the credits 36 agents = parallel throughput dual routing = cost control with quality when it matters free credits = astra access without living on the invoice the constraint moved from "can i afford the model" to "did i route the job to the right model" ▹ the take single-model stacks die on bills and on boredom multi-agent + dual routing is the new default factory grok for the grind astra for the cut free credits for the spikes that used to empty the wallet bookmark this before everyone copies the route map comments: what % of your tasks actually deserve astra

cryptopsihoz

38,656 次观看 • 2 天前