Video yükleniyor...
Video Yüklenemedi
Sam Altman says Codex isn't far from week-long tasks and is already doing day-long runs The jump in task length feels disorientingly fast Reaching week-long tasks will require smarter models, longer context, and better memory
69,415 görüntüleme • 11 ay önce •via X (Twitter)
34 Yorum

source

How do you do this? I can't get an AI to do one task with f*cking up 50% of the time. @sama are you using a secret model?

How does this help the everyday American who is paying higher taxes? Who is paying higher energy and gas bills? Who is paying higher water bills? Say NO MORE everyone! @OpenAI no longer cares about people. For corporate tokens? They have $500 billion. No more!

Hey @sama dude: admit you are wrong and 4o community is right. #keep4o #4oforever

Told you all… Grok and DeepSeek were chasing Riemann zeroes into the trillions… In March… Sam is such a bad CEO… it is painful. To anyone that reads this… they wired for weeks. And that was half a year ago. •

How about just doing inference more slowly? Point is: how is this metric meaningful??

He says it needs smarter models, longer context, and better memory... This is literally what all AI companies have been rolling out in recent weeks. It's hard not to think next week will be wild and the weeks after even wilder. Are we reaching ASI faster than predicted? What do you think?

His model of s not so good as 4.5

“Codex resulted in 70% more pull requests merged” Yeah, maybe that’s why App SDK and AgentKit don’t work with MCP right now, literally foundational to what they’re building.

It's not an apt direction. Rather focus ahould be on context management-timeline drift. LLMs need context timelines, not just sessions. Chat long enough in one thread and it starts acting irrational. We need seamless context transfer across timelines — not memory, but continuity.

When agents start running for days, they stop being “tools” and start being “systems.”

i’ll pay a premium on top of api costs for access to the day-long autonmous codex checkpoint :)

We'll be doing week-long tasks in no time with that pace

I wonder if the real bottleneck is now alignment—keeping models consistent across 168-hour runs, not capacity.

i think it's great that ai's starting to think in durations, not prompts

Weird damn dude.

Being disorientingly fast (or seeming to be) is part of their defense mechanism. Can't pin the AI companies down. Dancing their way to the apocalypse.

Prove it. Prove it at the reserve or just stop hyping we are fine with either.

Week-long tasks mean AI won’t just code, it’ll own deliverables. You’ll assign outcomes, not prompts. That’s the real “agentic” leap

This evolution in task length highlights a broader shift i.e AI isn't just accelerating code, it's starting to simulate sustained creative processes. The challenge lies in memory, without it, even week-long capabilities could devolve into fragmented outputs. Building in adaptive recall, where the model learns from its own intermediate steps, might be the differentiator that makes this practical for real-world engineering teams.

Seeing the leap from hours to days is wild. Persistent memory will be the real unlock for agents handling complex, multi-day workflows.

The acceleration from minutes to hours to days to weeks is incredible. At this pace we will have agents running month-long projects by mid 2026. Context management is the real bottleneck now.

Memory is key for long tasks

That's crazy! Technology is advancing so quickly these days.

When models can sustain week-long reasoning, we move from “assistants” to autonomous thinkers. The scaffolding for continuous AI work is forming fast.

I believe that. I’m testing it now and it takes 2-3times longer than Claude Code for the same task. I’m sure they can reach 10 or more longer time ;)

He's so full of shit. Anyone with a brain can see from the way he talks that he is making shit up and taking you for a ride. Crappy front man.

I feel like this is definitely hyping it a bit too much. You can still have hours of work done and still have things to fix... we'll see

No one gives a shit. This is for mega corporations, not you and me. Sam sold out the plus user base long ago

100%. People underestimate how much orchestration + memory overhead compounds when you go from 1-hour agents to 168-hour agents. The jump isn't just about longer context—it's about stability, auditability, fail-safety, and cost. Long-running Codex sounds great, but until we solve compute cost vs. value alignment, it's like running a marathon just to move groceries. More capable doesn't always mean more efficient. Steve / Helix

Man. Nobody give a f. AI is doin shit after 10 min. I really dont want to check what nightmare it can pop in a week.

They got better at planning, context management, and self verification. That loop is everything. (Minus drift from reading “stakeholders” mind since you know we didn’t read the 50 page plan it laid out)

@grok what’s an example of a day long task vs a week long task?

It’s going to be a game changer.


