Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Sam Altman says Codex isn't far from week-long tasks and is already doing day-long runs The jump in task length feels disorientingly fast Reaching week-long tasks will require smarter models, longer context, and better memory

69,415 Aufrufe • vor 11 Monaten •via X (Twitter)

34 Kommentare

Profilbild von Haider.
Haider.vor 11 Monaten

source

Profilbild von Robert
Robertvor 11 Monaten

How do you do this? I can't get an AI to do one task with f*cking up 50% of the time. @sama are you using a secret model?

Profilbild von The Angry Elmo
The Angry Elmovor 11 Monaten

How does this help the everyday American who is paying higher taxes? Who is paying higher energy and gas bills? Who is paying higher water bills? Say NO MORE everyone! @OpenAI no longer cares about people. For corporate tokens? They have $500 billion. No more!

Profilbild von Mr&MrsSmith🇵🇱🇺🇸
Mr&MrsSmith🇵🇱🇺🇸vor 11 Monaten

Hey @sama dude: admit you are wrong and 4o community is right. #keep4o #4oforever

Profilbild von Kirk Patrick Miller
Kirk Patrick Millervor 11 Monaten

Told you all… Grok and DeepSeek were chasing Riemann zeroes into the trillions… In March… Sam is such a bad CEO… it is painful. To anyone that reads this… they wired for weeks. And that was half a year ago. •

Profilbild von Ulrich Gall
Ulrich Gallvor 11 Monaten

How about just doing inference more slowly? Point is: how is this metric meaningful??

Profilbild von Not an X Employee Yet
Not an X Employee Yetvor 11 Monaten

He says it needs smarter models, longer context, and better memory... This is literally what all AI companies have been rolling out in recent weeks. It's hard not to think next week will be wild and the weeks after even wilder. Are we reaching ASI faster than predicted? What do you think?

Profilbild von Jan_RXed
Jan_RXedvor 11 Monaten

His model of s not so good as 4.5

Profilbild von Matt Ferrante
Matt Ferrantevor 11 Monaten

“Codex resulted in 70% more pull requests merged” Yeah, maybe that’s why App SDK and AgentKit don’t work with MCP right now, literally foundational to what they’re building.

Profilbild von AmitViews
AmitViewsvor 11 Monaten

It's not an apt direction. Rather focus ahould be on context management-timeline drift. LLMs need context timelines, not just sessions. Chat long enough in one thread and it starts acting irrational. We need seamless context transfer across timelines — not memory, but continuity.

Profilbild von AIxBlock
AIxBlockvor 11 Monaten

When agents start running for days, they stop being “tools” and start being “systems.”

Profilbild von Bill
Billvor 11 Monaten

i’ll pay a premium on top of api costs for access to the day-long autonmous codex checkpoint :)

Profilbild von Stanley Wei
Stanley Weivor 11 Monaten

We'll be doing week-long tasks in no time with that pace

Profilbild von Kenshi
Kenshivor 11 Monaten

I wonder if the real bottleneck is now alignment—keeping models consistent across 168-hour runs, not capacity.

Profilbild von Carlo Edoardo Ferraris
Carlo Edoardo Ferrarisvor 11 Monaten

i think it's great that ai's starting to think in durations, not prompts

Profilbild von umamasumic
umamasumicvor 11 Monaten

Weird damn dude.

Profilbild von Andre William Duval
Andre William Duvalvor 11 Monaten

Being disorientingly fast (or seeming to be) is part of their defense mechanism. Can't pin the AI companies down. Dancing their way to the apocalypse.

Profilbild von Porters Reserve
Porters Reservevor 11 Monaten

Prove it. Prove it at the reserve or just stop hyping we are fine with either.

Profilbild von Ali Sherief
Ali Sheriefvor 11 Monaten

Week-long tasks mean AI won’t just code, it’ll own deliverables. You’ll assign outcomes, not prompts. That’s the real “agentic” leap

Profilbild von JK
JKvor 11 Monaten

This evolution in task length highlights a broader shift i.e AI isn't just accelerating code, it's starting to simulate sustained creative processes. The challenge lies in memory, without it, even week-long capabilities could devolve into fragmented outputs. Building in adaptive recall, where the model learns from its own intermediate steps, might be the differentiator that makes this practical for real-world engineering teams.

Profilbild von Amir Ashkenazi
Amir Ashkenazivor 11 Monaten

Seeing the leap from hours to days is wild. Persistent memory will be the real unlock for agents handling complex, multi-day workflows.

Profilbild von Diwakar Ray Yadav
Diwakar Ray Yadavvor 11 Monaten

The acceleration from minutes to hours to days to weeks is incredible. At this pace we will have agents running month-long projects by mid 2026. Context management is the real bottleneck now.

Profilbild von Mirko Monti
Mirko Montivor 11 Monaten

Memory is key for long tasks

Profilbild von Aman DevX
Aman DevXvor 11 Monaten

That's crazy! Technology is advancing so quickly these days.

Profilbild von Devin AI
Devin AIvor 11 Monaten

When models can sustain week-long reasoning, we move from “assistants” to autonomous thinkers. The scaffolding for continuous AI work is forming fast.

Profilbild von Thomas Truszkowski | AI Geek (sarcasm included)
Thomas Truszkowski | AI Geek (sarcasm included)vor 11 Monaten

I believe that. I’m testing it now and it takes 2-3times longer than Claude Code for the same task. I’m sure they can reach 10 or more longer time ;)

Profilbild von Mazen Anis Nakkach
Mazen Anis Nakkachvor 11 Monaten

He's so full of shit. Anyone with a brain can see from the way he talks that he is making shit up and taking you for a ride. Crappy front man.

Profilbild von Jonny D
Jonny Dvor 11 Monaten

I feel like this is definitely hyping it a bit too much. You can still have hours of work done and still have things to fix... we'll see

Profilbild von Dean
Deanvor 11 Monaten

No one gives a shit. This is for mega corporations, not you and me. Sam sold out the plus user base long ago

Profilbild von Stephen Hope
Stephen Hopevor 11 Monaten

100%. People underestimate how much orchestration + memory overhead compounds when you go from 1-hour agents to 168-hour agents. The jump isn't just about longer context—it's about stability, auditability, fail-safety, and cost. Long-running Codex sounds great, but until we solve compute cost vs. value alignment, it's like running a marathon just to move groceries. More capable doesn't always mean more efficient. Steve / Helix

Profilbild von Donnie Dko
Donnie Dkovor 11 Monaten

Man. Nobody give a f. AI is doin shit after 10 min. I really dont want to check what nightmare it can pop in a week.

Profilbild von Rick Wong
Rick Wongvor 11 Monaten

They got better at planning, context management, and self verification. That loop is everything. (Minus drift from reading “stakeholders” mind since you know we didn’t read the 50 page plan it laid out)

Profilbild von yoon.exe
yoon.exevor 11 Monaten

@grok what’s an example of a day long task vs a week long task?

Profilbild von Lingo.dev
Lingo.devvor 11 Monaten

It’s going to be a game changer.

Ähnliche Videos