Loading video...

Video Failed to Load

Go Home

Sam Altman says Codex isn't far from week-long tasks and is already doing day-long runs The jump in task length feels disorientingly fast Reaching week-long tasks will require smarter models, longer context, and better memory

69,415 views • 11 months ago •via X (Twitter)

34 Comments

Haider.'s profile picture
Haider.11 months ago

source

Robert's profile picture
Robert11 months ago

How do you do this? I can't get an AI to do one task with f*cking up 50% of the time. @sama are you using a secret model?

The Angry Elmo's profile picture
The Angry Elmo11 months ago

How does this help the everyday American who is paying higher taxes? Who is paying higher energy and gas bills? Who is paying higher water bills? Say NO MORE everyone! @OpenAI no longer cares about people. For corporate tokens? They have $500 billion. No more!

Mr&MrsSmith🇵🇱🇺🇸's profile picture
Mr&MrsSmith🇵🇱🇺🇸11 months ago

Hey @sama dude: admit you are wrong and 4o community is right. #keep4o #4oforever

Kirk Patrick Miller's profile picture
Kirk Patrick Miller11 months ago

Told you all… Grok and DeepSeek were chasing Riemann zeroes into the trillions… In March… Sam is such a bad CEO… it is painful. To anyone that reads this… they wired for weeks. And that was half a year ago. •

Ulrich Gall's profile picture
Ulrich Gall11 months ago

How about just doing inference more slowly? Point is: how is this metric meaningful??

Not an X Employee Yet's profile picture
Not an X Employee Yet11 months ago

He says it needs smarter models, longer context, and better memory... This is literally what all AI companies have been rolling out in recent weeks. It's hard not to think next week will be wild and the weeks after even wilder. Are we reaching ASI faster than predicted? What do you think?

Jan_RXed's profile picture
Jan_RXed11 months ago

His model of s not so good as 4.5

Matt Ferrante's profile picture
Matt Ferrante11 months ago

“Codex resulted in 70% more pull requests merged” Yeah, maybe that’s why App SDK and AgentKit don’t work with MCP right now, literally foundational to what they’re building.

AmitViews's profile picture
AmitViews11 months ago

It's not an apt direction. Rather focus ahould be on context management-timeline drift. LLMs need context timelines, not just sessions. Chat long enough in one thread and it starts acting irrational. We need seamless context transfer across timelines — not memory, but continuity.

AIxBlock's profile picture
AIxBlock11 months ago

When agents start running for days, they stop being “tools” and start being “systems.”

Bill's profile picture
Bill11 months ago

i’ll pay a premium on top of api costs for access to the day-long autonmous codex checkpoint :)

Stanley Wei's profile picture
Stanley Wei11 months ago

We'll be doing week-long tasks in no time with that pace

Kenshi's profile picture
Kenshi11 months ago

I wonder if the real bottleneck is now alignment—keeping models consistent across 168-hour runs, not capacity.

Carlo Edoardo Ferraris's profile picture
Carlo Edoardo Ferraris11 months ago

i think it's great that ai's starting to think in durations, not prompts

umamasumic's profile picture
umamasumic11 months ago

Weird damn dude.

Andre William Duval's profile picture
Andre William Duval11 months ago

Being disorientingly fast (or seeming to be) is part of their defense mechanism. Can't pin the AI companies down. Dancing their way to the apocalypse.

Porters Reserve's profile picture
Porters Reserve11 months ago

Prove it. Prove it at the reserve or just stop hyping we are fine with either.

Ali Sherief's profile picture
Ali Sherief11 months ago

Week-long tasks mean AI won’t just code, it’ll own deliverables. You’ll assign outcomes, not prompts. That’s the real “agentic” leap

JK's profile picture
JK11 months ago

This evolution in task length highlights a broader shift i.e AI isn't just accelerating code, it's starting to simulate sustained creative processes. The challenge lies in memory, without it, even week-long capabilities could devolve into fragmented outputs. Building in adaptive recall, where the model learns from its own intermediate steps, might be the differentiator that makes this practical for real-world engineering teams.

Amir Ashkenazi's profile picture
Amir Ashkenazi11 months ago

Seeing the leap from hours to days is wild. Persistent memory will be the real unlock for agents handling complex, multi-day workflows.

Diwakar Ray Yadav's profile picture
Diwakar Ray Yadav11 months ago

The acceleration from minutes to hours to days to weeks is incredible. At this pace we will have agents running month-long projects by mid 2026. Context management is the real bottleneck now.

Mirko Monti's profile picture
Mirko Monti11 months ago

Memory is key for long tasks

Aman DevX's profile picture
Aman DevX11 months ago

That's crazy! Technology is advancing so quickly these days.

Devin AI's profile picture
Devin AI11 months ago

When models can sustain week-long reasoning, we move from “assistants” to autonomous thinkers. The scaffolding for continuous AI work is forming fast.

Thomas Truszkowski | AI Geek (sarcasm included)'s profile picture
Thomas Truszkowski | AI Geek (sarcasm included)11 months ago

I believe that. I’m testing it now and it takes 2-3times longer than Claude Code for the same task. I’m sure they can reach 10 or more longer time ;)

Mazen Anis Nakkach's profile picture
Mazen Anis Nakkach11 months ago

He's so full of shit. Anyone with a brain can see from the way he talks that he is making shit up and taking you for a ride. Crappy front man.

Jonny D's profile picture
Jonny D11 months ago

I feel like this is definitely hyping it a bit too much. You can still have hours of work done and still have things to fix... we'll see

Dean's profile picture
Dean11 months ago

No one gives a shit. This is for mega corporations, not you and me. Sam sold out the plus user base long ago

Stephen Hope's profile picture
Stephen Hope11 months ago

100%. People underestimate how much orchestration + memory overhead compounds when you go from 1-hour agents to 168-hour agents. The jump isn't just about longer context—it's about stability, auditability, fail-safety, and cost. Long-running Codex sounds great, but until we solve compute cost vs. value alignment, it's like running a marathon just to move groceries. More capable doesn't always mean more efficient. Steve / Helix

Donnie Dko's profile picture
Donnie Dko11 months ago

Man. Nobody give a f. AI is doin shit after 10 min. I really dont want to check what nightmare it can pop in a week.

Rick Wong's profile picture
Rick Wong11 months ago

They got better at planning, context management, and self verification. That loop is everything. (Minus drift from reading “stakeholders” mind since you know we didn’t read the 50 page plan it laid out)

yoon.exe's profile picture
yoon.exe11 months ago

@grok what’s an example of a day long task vs a week long task?

Lingo.dev's profile picture
Lingo.dev11 months ago

It’s going to be a game changer.

Related Videos