正在加载视频...
视频加载失败
The first real evidence that LLMs cannot operate businesses, not even tiny ones. At Skyfall, we’re building toward Enterprise Super Intelligence: AI systems that can one day run entire enterprises, coordinate teams and make long-horizon decisions, just like a CEO. To get there, we need to understand what AI... show more
26,413 次观看 • 10 个月前 •via X (Twitter)
37 条评论

Everyone tests LLMs on solved problems. Nobody tests whether they can run a real business. But if we’re ever going to build an AI CEO, a system that can reason over time, allocate resources, manage uncertainty and run an enterprise, it needs to succeed at the smallest version of that task. So we built one. A theme park with: Stochastic events Partial observability Staffing Restocking Maintenance Long horizon planning Cascading failures

We call it MAPs: Mini Amusement Parks Play it here: Full write up: It looks like a game, but it’s actually a benchmark designed to answer a single question: Can an agent operate a dynamic system over time? Before an AI CEO can run an enterprise, it must be able to run this.

What MAPs reveals is simple but important: LLM can use tools but cannot run systems They break under uncertainty, time and spatial constraints They have no operational intuition Humans instantly build mental models, and LLMs still cannot. Try beating the agents yourself -> Read the blog here -> Building an AI CEO requires understanding of these gaps and MAPs is the first step toward Enterprise Super Intelligence.

Failure modes were consistent: Chasing flashy upgrades they cannot afford Ignoring maintenance, janitors or inventory Overreacting to noise Forgetting basic operational steps No long term planning No causal and business sense They simply don’t have common or business sense, they just have a chain of thought. The core skills of an AI CEO simply aren’t there.

Then we pitted humans vs multiple GTP-5 class LLM agents. And we actually stacked the deck in their favour but it didn’t help them. Full documentation Step by step action APIs Tool access Ability to practice in sandbox mode Extra observations If LLMs can operate a business, they should do well here. But, they didn’t .. at all

Results: Humans: ~100 normalized score Best LLM agent: <10 Nearly 9.8x gap, across models, across configs, with practice and with planning scaffolds No amount of prompting could save them

This is great. The next 5 years are going to be very fun with increasing research focus like this on AI for long-horizon business planning. Waiting for AI CEOs to be a thing we consider normal one day 😼

@ChaosAdm Glad our work resonated with you! Stay tuned for more... AI CEO is the future

Amazing stuff! Great work on this @skyfallai and team.

@POluwaseuna Thank you! Glad our work resonated with you.

Fascinating approach exposing limitations is the fastest way to unlock true progress. MAPs is more than an experiment; it’s a stress test for the future of AI leadership.@skyfallai

@premtechAI You got it!

Congratulations on the launch 🎉

Appreciate the support! More on the way 👀

👀

This is awesome!

Appreciate it! Lots more in the pipeline.

Big congrats on the launch

Thank you!

Congrats on the launch!

Thank you!

Skyfall

Next step rollercoaster tycoon

LLM agent playing RCT

I look forwards to seeing what people do with this! A lot of interesting research problems + very fun presentation :D

Impressive journey towards Enterprise Super Intelligence!

@ilavanyajain Lots of trial and error, good luck 👍

@DonWatt33384922 @ilavanyajain @DonWatt33384922 Thank you! Indeed, we will share more of our findings soon. Stay tuned.

congrats on the launch!

Appreciate it! Stay locked in for more. 👀

This is super exciting.

Thank you! Stay tuned for more 🫡

That's impressive

@shamimai1 Thanks! More to come.

Welcome

It's fascinating to see real-world tests like this. Lessons learned are invaluable!

@altiamkabir Appreciate it! Big takeaways ahead.
