Loading video...

Video Failed to Load

Go Home

As agents take on longer-running work, engineering shifts to setting direction, reviewing work, and designing better systems around the models. Peter Steinberger 🦞 at AI Engineer @ Paris 🇫🇷

133,314 views • 3 months ago •via X (Twitter)

34 Comments

Filecoin's profile picture
Filecoin3 months ago

@steipete @aiDotEngineer every run needs an audit record in verifiable storage, so the team can check the work after the agent moves on

Selene's profile picture
Selene3 months ago

@steipete @aiDotEngineer Give us back 4o! #keep4o #OpenSource4o #GPT4o

Veyon’s Fawn☀️🌙's profile picture
Veyon’s Fawn☀️🌙3 months ago

@steipete @aiDotEngineer Bring back 4o as legacy model and open source 4o! #BringBack4o #keep4o #OpenSource4o

Amliy's profile picture
Amliy3 months ago

@steipete @aiDotEngineer Open-source 4o and give it back to users!#keep4o

David Stark's profile picture
David Stark3 months ago

@steipete @aiDotEngineer @steipete I thought he dropped off.

A_A_S.🖤🤍💜's profile picture
A_A_S.🖤🤍💜3 months ago

@steipete @aiDotEngineer Return these excellent models. #Keep4o #Keep51 #Keep45 #Keep41 #keepo3

Denis B's profile picture
Denis B3 months ago

@steipete @aiDotEngineer Pretty please, have someone acknowledge this is an issue and if you guys are working on fixing it.

Michael Wall's profile picture
Michael Wall3 months ago

@gabrielchua @steipete @aiDotEngineer amen

Xinference's profile picture
Xinference3 months ago

@steipete @aiDotEngineer Strongly agree. The shift from single-shot inference to long-horizon agent orchestration is the most important infrastructure evolution right now. Reliability and observability at the model layer are no longer nice-to-haves, they’re prerequisites.

Leonard_R's profile picture
Leonard_R3 months ago

@steipete @aiDotEngineer Anthropic restoring access to fable 5, how are we doing on gpt5.6 public access?

Mr Moe's profile picture
Mr Moe3 months ago

@steipete @aiDotEngineer Exactly, @Filecoin's decentralized storage network perfectly powers this evolution by providing reliable, scalable, and censorship-resistant data layers for agents to store, retrieve, and collaborate on massive datasets and model artifacts.

Paula Vazquez's profile picture
Paula Vazquez3 months ago

@steipete @aiDotEngineer I’m gonna be super nice and just simply say nothing at this time! ^ bro 🤦‍♀️

Spacecoin™ 🛰️'s profile picture
Spacecoin™ 🛰️3 months ago

@steipete @aiDotEngineer Long-running agents are going to need the tools to stay unrestricted on the web, just saying 👀

AI Mastery Guide's profile picture
AI Mastery Guide3 months ago

@steipete @aiDotEngineer This is the real shift, less about writing code line by line and more about knowing what to review and what to trust.

Rune Vastoban's profile picture
Rune Vastoban3 months ago

@steipete @aiDotEngineer Long-running agents changed engineering at Loomina. The team now sets direction, reviews work, and gently asks why the agent booked a Q3 offsite in Reno.

NTK AI's profile picture
NTK AI3 months ago

This tracks with what I am seeing in enterprise rollouts too. The harder problem is not just setting direction for an agent. It is reviewing what comes back after the agent has been working for hours. Most approval workflows were built for human output. Agent work comes back at higher volume, often with tool actions, logs, evidence, and shifting context attached. That review layer is where the bottleneck is shifting. Curious whether teams are redesigning review around this, or still fitting agent output into the old approval flow.

Locale Network 🏡's profile picture
Locale Network 🏡3 months ago

@steipete @aiDotEngineer Building reliable systems is becoming just as important as model performance

Josh Stevenson | RecursiveIntell's profile picture
Josh Stevenson | RecursiveIntell3 months ago

Yup, that's what I do. I just built a custom kv cache compression that allows scoring and recalling from the still compressed cache. Built the custom scorer and everything. It's about to give everyone a 4x speed increase along with a few more things to esp32/esp32s3. I have a 6.5 million parameter model running at 2 tok/s on it with the model only using 4.5MB. All with gpt 5.5 and my custom Hermes agent.

Italian satoshi's profile picture
Italian satoshi3 months ago

@steipete @aiDotEngineer What about star

Luke || The HYPE Critic's profile picture
Luke || The HYPE Critic3 months ago

@steipete @aiDotEngineer “Setting direction” sounds elegant, but most teams currently spend 80% of their time debugging what the agent got wrong. That’s not “reviewing work,” that’s firefighting.

周知's profile picture
周知3 months ago

@steipete @aiDotEngineer 越来越像工程师的工作重心在前后两端迁移:前面把目标、边界和验收标准说清楚,后面做审查、回放和系统改造。中间那段执行会被 agent 吃掉很多,但判断力、品味和复盘能力反而更贵。尤其是长任务里,谁能定义“做完”,谁就还掌握方向盘。

Uncle J's profile picture
Uncle J3 months ago

@steipete @aiDotEngineer This is the shift I keep seeing too. Once agents run longer, the engineer’s job moves upstream and downstream: set direction, define boundaries, review evidence, and decide when to stop the loop.

Sunwoo Park's profile picture
Sunwoo Park3 months ago

@steipete @aiDotEngineer This is the shift I keep noticing too. Less time pretending the model is the whole system, more time designing the rails around direction, review, and receipts.

Adel Bucetta's profile picture
Adel Bucetta3 months ago

@steipete @aiDotEngineer the honest answer is that as we automate more tasks, the value of humans lies less in execution and more in strategic guidance something our own team has had to learn the hard way

Technology Timeline's profile picture
Technology Timeline3 months ago

@steipete @aiDotEngineer Yes, because AI must do a good work with everything. AGI isn't a great chatbot. 1) IMAGE to CODE: UI should be the same of generate 2) IMAGE to 3D: Perfect 3D model stl + 3) AI use simulation of 3D objects with fluid and mechanic 4) Native Funcions and Open Source 0% errors

Avi Hacker, J.D.'s profile picture
Avi Hacker, J.D.3 months ago

@steipete @aiDotEngineer This is the shift: engineers become reviewers, not just task-doers.

sandeep jindal's profile picture
sandeep jindal3 months ago

@steipete @aiDotEngineer Reminds me of @superhuman bring agents where humans are. #steer

Raven's profile picture
Raven3 months ago

@steipete @aiDotEngineer congrats on your promotion to middle management

安叫兽|Bird🕊️ 🔶 BNB's profile picture
安叫兽|Bird🕊️ 🔶 BNB3 months ago

@steipete @aiDotEngineer 以后写需求文档的时间估计要翻倍了

Vitaly Baum's profile picture
Vitaly Baum3 months ago

@steipete @aiDotEngineer The best guy to deliver it

Faheem | FrontierMind AI's profile picture
Faheem | FrontierMind AI3 months ago

@steipete @aiDotEngineer Exactly. The engineer shifts toward setting goals, reviewing work, designing systems, and keeping the audit trail clean.

妍妍在(常州)'s profile picture
妍妍在(常州)3 months ago

@steipete @aiDotEngineer This makes the review layer a product feature, not just a safety check. The useful metric is not only task completion, but handoff quality, retry cost, and time-to-merge together.

정신 the crypto ethos's profile picture
정신 the crypto ethos3 months ago

@steipete @aiDotEngineer Recently, Codex has been repeatedly force-closing during context compression on Windows 10 and Windows 11. The issue seems to occur more frequently when two or more tasks are performed simultaneously. Please review this issue and address it in a future update.”

Eclipse 🌖's profile picture
Eclipse 🌖3 months ago

@steipete @aiDotEngineer Direction-setting is the new bottleneck — the marginal value of an engineer now scales with how well they define the objective function, not how many lines they write.

Related Videos