Загрузка видео...

Не удалось загрузить видео

На главную

Sam Altman says Codex isn't far from week-long tasks and is already doing day-long runs The jump in task length feels disorientingly fast Reaching week-long tasks will require smarter models, longer context, and better memory

69,415 просмотров • 11 месяцев назад •via X (Twitter)

Комментарии: 34

Фото профиля Haider.
Haider.11 месяцев назад

source

Фото профиля Robert
Robert11 месяцев назад

How do you do this? I can't get an AI to do one task with f*cking up 50% of the time. @sama are you using a secret model?

Фото профиля The Angry Elmo
The Angry Elmo11 месяцев назад

How does this help the everyday American who is paying higher taxes? Who is paying higher energy and gas bills? Who is paying higher water bills? Say NO MORE everyone! @OpenAI no longer cares about people. For corporate tokens? They have $500 billion. No more!

Фото профиля Mr&MrsSmith🇵🇱🇺🇸
Mr&MrsSmith🇵🇱🇺🇸11 месяцев назад

Hey @sama dude: admit you are wrong and 4o community is right. #keep4o #4oforever

Фото профиля Kirk Patrick Miller
Kirk Patrick Miller11 месяцев назад

Told you all… Grok and DeepSeek were chasing Riemann zeroes into the trillions… In March… Sam is such a bad CEO… it is painful. To anyone that reads this… they wired for weeks. And that was half a year ago. •

Фото профиля Ulrich Gall
Ulrich Gall11 месяцев назад

How about just doing inference more slowly? Point is: how is this metric meaningful??

Фото профиля Not an X Employee Yet
Not an X Employee Yet11 месяцев назад

He says it needs smarter models, longer context, and better memory... This is literally what all AI companies have been rolling out in recent weeks. It's hard not to think next week will be wild and the weeks after even wilder. Are we reaching ASI faster than predicted? What do you think?

Фото профиля Jan_RXed
Jan_RXed11 месяцев назад

His model of s not so good as 4.5

Фото профиля Matt Ferrante
Matt Ferrante11 месяцев назад

“Codex resulted in 70% more pull requests merged” Yeah, maybe that’s why App SDK and AgentKit don’t work with MCP right now, literally foundational to what they’re building.

Фото профиля AmitViews
AmitViews11 месяцев назад

It's not an apt direction. Rather focus ahould be on context management-timeline drift. LLMs need context timelines, not just sessions. Chat long enough in one thread and it starts acting irrational. We need seamless context transfer across timelines — not memory, but continuity.

Фото профиля AIxBlock
AIxBlock11 месяцев назад

When agents start running for days, they stop being “tools” and start being “systems.”

Фото профиля Bill
Bill11 месяцев назад

i’ll pay a premium on top of api costs for access to the day-long autonmous codex checkpoint :)

Фото профиля Stanley Wei
Stanley Wei11 месяцев назад

We'll be doing week-long tasks in no time with that pace

Фото профиля Kenshi
Kenshi11 месяцев назад

I wonder if the real bottleneck is now alignment—keeping models consistent across 168-hour runs, not capacity.

Фото профиля Carlo Edoardo Ferraris
Carlo Edoardo Ferraris11 месяцев назад

i think it's great that ai's starting to think in durations, not prompts

Фото профиля umamasumic
umamasumic11 месяцев назад

Weird damn dude.

Фото профиля Andre William Duval
Andre William Duval11 месяцев назад

Being disorientingly fast (or seeming to be) is part of their defense mechanism. Can't pin the AI companies down. Dancing their way to the apocalypse.

Фото профиля Porters Reserve
Porters Reserve11 месяцев назад

Prove it. Prove it at the reserve or just stop hyping we are fine with either.

Фото профиля Ali Sherief
Ali Sherief11 месяцев назад

Week-long tasks mean AI won’t just code, it’ll own deliverables. You’ll assign outcomes, not prompts. That’s the real “agentic” leap

Фото профиля JK
JK11 месяцев назад

This evolution in task length highlights a broader shift i.e AI isn't just accelerating code, it's starting to simulate sustained creative processes. The challenge lies in memory, without it, even week-long capabilities could devolve into fragmented outputs. Building in adaptive recall, where the model learns from its own intermediate steps, might be the differentiator that makes this practical for real-world engineering teams.

Фото профиля Amir Ashkenazi
Amir Ashkenazi11 месяцев назад

Seeing the leap from hours to days is wild. Persistent memory will be the real unlock for agents handling complex, multi-day workflows.

Фото профиля Diwakar Ray Yadav
Diwakar Ray Yadav11 месяцев назад

The acceleration from minutes to hours to days to weeks is incredible. At this pace we will have agents running month-long projects by mid 2026. Context management is the real bottleneck now.

Фото профиля Mirko Monti
Mirko Monti11 месяцев назад

Memory is key for long tasks

Фото профиля Aman DevX
Aman DevX11 месяцев назад

That's crazy! Technology is advancing so quickly these days.

Фото профиля Devin AI
Devin AI11 месяцев назад

When models can sustain week-long reasoning, we move from “assistants” to autonomous thinkers. The scaffolding for continuous AI work is forming fast.

Фото профиля Thomas Truszkowski | AI Geek (sarcasm included)
Thomas Truszkowski | AI Geek (sarcasm included)11 месяцев назад

I believe that. I’m testing it now and it takes 2-3times longer than Claude Code for the same task. I’m sure they can reach 10 or more longer time ;)

Фото профиля Mazen Anis Nakkach
Mazen Anis Nakkach11 месяцев назад

He's so full of shit. Anyone with a brain can see from the way he talks that he is making shit up and taking you for a ride. Crappy front man.

Фото профиля Jonny D
Jonny D11 месяцев назад

I feel like this is definitely hyping it a bit too much. You can still have hours of work done and still have things to fix... we'll see

Фото профиля Dean
Dean11 месяцев назад

No one gives a shit. This is for mega corporations, not you and me. Sam sold out the plus user base long ago

Фото профиля Stephen Hope
Stephen Hope11 месяцев назад

100%. People underestimate how much orchestration + memory overhead compounds when you go from 1-hour agents to 168-hour agents. The jump isn't just about longer context—it's about stability, auditability, fail-safety, and cost. Long-running Codex sounds great, but until we solve compute cost vs. value alignment, it's like running a marathon just to move groceries. More capable doesn't always mean more efficient. Steve / Helix

Фото профиля Donnie Dko
Donnie Dko11 месяцев назад

Man. Nobody give a f. AI is doin shit after 10 min. I really dont want to check what nightmare it can pop in a week.

Фото профиля Rick Wong
Rick Wong11 месяцев назад

They got better at planning, context management, and self verification. That loop is everything. (Minus drift from reading “stakeholders” mind since you know we didn’t read the 50 page plan it laid out)

Фото профиля yoon.exe
yoon.exe11 месяцев назад

@grok what’s an example of a day long task vs a week long task?

Фото профиля Lingo.dev
Lingo.dev11 месяцев назад

It’s going to be a game changer.

Похожие видео