正在加载视频...

视频加载失败

The first real evidence that LLMs cannot operate businesses, not even tiny ones. At Skyfall, we’re building toward Enterprise Super Intelligence: AI systems that can one day run entire enterprises, coordinate teams and make long-horizon decisions, just like a CEO. To get there, we need to understand what AI...

26,413 次观看 • 10 个月前 •via X (Twitter)

37 条评论

Skyfall AI 的头像
Skyfall AI10 个月前

Everyone tests LLMs on solved problems. Nobody tests whether they can run a real business. But if we’re ever going to build an AI CEO, a system that can reason over time, allocate resources, manage uncertainty and run an enterprise, it needs to succeed at the smallest version of that task. So we built one. A theme park with: Stochastic events Partial observability Staffing Restocking Maintenance Long horizon planning Cascading failures

Skyfall AI 的头像
Skyfall AI10 个月前

We call it MAPs: Mini Amusement Parks Play it here: Full write up: It looks like a game, but it’s actually a benchmark designed to answer a single question: Can an agent operate a dynamic system over time? Before an AI CEO can run an enterprise, it must be able to run this.

Skyfall AI 的头像
Skyfall AI10 个月前

What MAPs reveals is simple but important: LLM can use tools but cannot run systems They break under uncertainty, time and spatial constraints They have no operational intuition Humans instantly build mental models, and LLMs still cannot. Try beating the agents yourself -> Read the blog here -> Building an AI CEO requires understanding of these gaps and MAPs is the first step toward Enterprise Super Intelligence.

Skyfall AI 的头像
Skyfall AI10 个月前

Failure modes were consistent: Chasing flashy upgrades they cannot afford Ignoring maintenance, janitors or inventory Overreacting to noise Forgetting basic operational steps No long term planning No causal and business sense They simply don’t have common or business sense, they just have a chain of thought. The core skills of an AI CEO simply aren’t there.

Skyfall AI 的头像
Skyfall AI10 个月前

Then we pitted humans vs multiple GTP-5 class LLM agents. And we actually stacked the deck in their favour but it didn’t help them. Full documentation Step by step action APIs Tool access Ability to practice in sandbox mode Extra observations If LLMs can operate a business, they should do well here. But, they didn’t .. at all

Skyfall AI 的头像
Skyfall AI10 个月前

Results: Humans: ~100 normalized score Best LLM agent: <10 Nearly 9.8x gap, across models, across configs, with practice and with planning scaffolds No amount of prompting could save them

Lakshya Gupta 的头像
Lakshya Gupta10 个月前

This is great. The next 5 years are going to be very fun with increasing research focus like this on AI for long-horizon business planning. Waiting for AI CEOs to be a thing we consider normal one day 😼

Skyfall AI 的头像
Skyfall AI10 个月前

@ChaosAdm Glad our work resonated with you! Stay tuned for more... AI CEO is the future

Oluwaseun | The Automation Guy - TAG 的头像
Oluwaseun | The Automation Guy - TAG10 个月前

Amazing stuff! Great work on this @skyfallai and team.

Skyfall AI 的头像
Skyfall AI10 个月前

@POluwaseuna Thank you! Glad our work resonated with you.

Prem 的头像
Prem10 个月前

Fascinating approach exposing limitations is the fastest way to unlock true progress. MAPs is more than an experiment; it’s a stress test for the future of AI leadership.@skyfallai

Skyfall AI 的头像
Skyfall AI10 个月前

@premtechAI You got it!

Flora_The_AI_Girl 的头像
Flora_The_AI_Girl10 个月前

Congratulations on the launch 🎉

Skyfall AI 的头像
Skyfall AI10 个月前

Appreciate the support! More on the way 👀

Dada 的头像
Dada10 个月前

👀

Brady Long 的头像
Brady Long10 个月前

This is awesome!

Skyfall AI 的头像
Skyfall AI10 个月前

Appreciate it! Lots more in the pipeline.

Alamin 的头像
Alamin10 个月前

Big congrats on the launch

Skyfall AI 的头像
Skyfall AI10 个月前

Thank you!

Nico Lecomte 的头像
Nico Lecomte10 个月前

Congrats on the launch!

Skyfall AI 的头像
Skyfall AI10 个月前

Thank you!

Spencer Baggins 的头像
Spencer Baggins10 个月前

Skyfall

Sai 的头像
Sai10 个月前

Next step rollercoaster tycoon

Skyfall AI 的头像
Skyfall AI10 个月前

LLM agent playing RCT

Ian Berlot-Attwell 的头像
Ian Berlot-Attwell10 个月前

I look forwards to seeing what people do with this! A lot of interesting research problems + very fun presentation :D

AI PlanetX 的头像
AI PlanetX10 个月前

Impressive journey towards Enterprise Super Intelligence!

Donald Watts 的头像
Donald Watts10 个月前

@ilavanyajain Lots of trial and error, good luck 👍

Skyfall AI 的头像
Skyfall AI10 个月前

@DonWatt33384922 @ilavanyajain @DonWatt33384922 Thank you! Indeed, we will share more of our findings soon. Stay tuned.

jpark 的头像
jpark10 个月前

congrats on the launch!

Skyfall AI 的头像
Skyfall AI10 个月前

Appreciate it! Stay locked in for more. 👀

Manu Ebert 的头像
Manu Ebert10 个月前

This is super exciting.

Skyfall AI 的头像
Skyfall AI10 个月前

Thank you! Stay tuned for more 🫡

Shamim Hossain 的头像
Shamim Hossain10 个月前

That's impressive

Skyfall AI 的头像
Skyfall AI10 个月前

@shamimai1 Thanks! More to come.

Shamim Hossain 的头像
Shamim Hossain10 个月前

Welcome

Altiam Kabir 的头像
Altiam Kabir10 个月前

It's fascinating to see real-world tests like this. Lessons learned are invaluable!

Skyfall AI 的头像
Skyfall AI10 个月前

@altiamkabir Appreciate it! Big takeaways ahead.

相关视频

Mansa AI is an enterprise-grade AI + Web3 platform designed to move artificial intelligence from experimentation into real-world execution. Built for creators, developers, and businesses, it focuses on deploying AI that actually works across modern digital systems, not just in isolated demos. 🚀 Production-ready AI infrastructure Mansa AI enables teams to deploy AI systems designed for live environments, handling real workflows, real data, and real operational demands without constant manual oversight. 🧠 Autonomous AI agents At its core, Mansa AI allows users to build autonomous agents that automate decision-making, coordinate tasks, monitor live signals, and execute complex workflows across dynamic environments. ⚙️ Fully customizable logic Agents can be configured with custom behaviors, triggers, and responses. From content generation and analytics to operational automation and intelligent orchestration, logic adapts to specific business strategies. 🔗 Web3 and off-chain integration Mansa AI bridges blockchain ecosystems with traditional systems, enabling cross-chain coordination, smart contract interactions, and seamless integration with existing enterprise infrastructure. 📊 Real-world use cases The platform supports automation for operations, customer engagement, analytics, data pipelines, content workflows, and AI-driven optimization across products and teams. 📈 Built for scale Whether launching as a startup or deploying across enterprise systems, Mansa AI is designed to scale AI operations without adding complexity or fragmentation. Mansa AI transforms artificial intelligence into deployable infrastructure. By combining autonomy, customization, interoperability, and scalability, it enables teams to own, operate, and grow intelligent systems that deliver real value in production environments.

King

155,637 次观看 • 9 个月前

Terence Tao has an IQ above 200. Youngest gold medalist in Math Olympiad history. Fields Medal winner. The greatest living mathematician by nearly any measure. And he just said something most people aren’t ready for. Tao: “This whole era of AI is teaching us that our idea of what intelligence is, is not really accurate.” We spent centuries building civilization on one assumption. That intelligence was sacred. Irreducible. Uniquely ours. The one thing that made the entire human story make sense. Then AI started solving things we swore only we could. Chess. Language. Vision. Math. And every time, we reached for the same defense. That’s not real intelligence. It’s just tricks. Just pattern matching. Just an algorithm. Tao: “You look at how it’s done and it doesn’t feel like intelligence.” So we moved the line. Again. And again. And again. Because intelligence was supposed to feel like something. Something deep. Something we could point to and say… this is what separates us from everything else. But AI kept solving the problems. And that feeling never arrived. Tao: “We were looking for some elusive, intelligent way of thinking and we don’t see it in the tools that actually solve our goals.” Here’s what makes it worse. Large language models work by predicting the next word. One word at a time. No grand architecture. No deep understanding. Just probability. And it works. Tao: “Maybe that’s actually a lot of what humans do as well.” The greatest living mathematician just told you human thought might run on the same machinery. Not some transcendent spark. Pattern recognition. Prediction. One thought, one decision, one word at a time. We built religion around intelligence. Philosophy around it. An entire species identity around it. And a machine running probability just held up a mirror. We didn’t lose intelligence to AI. We just finally saw what it always was. What haunts us isn’t that machines learned to think. It’s that thinking was never what we needed it to be.

Dustin

565,514 次观看 • 4 个月前

Elon Musk just worked out where humanity actually fits in the future, and it’s smaller than you think. Biological intelligence ends up being less than 1% of all intelligence that will exist. The rest? Artificial. Musk: “The vast majority of future intelligence will be AI.” We’re the match, not the bonfire. Which makes alignment the problem that determines whether we survive or disappear. Get it right and that 99% helps us reach the stars. Get it wrong and we’re just temporary infrastructure that got replaced the second something smarter came along. We’re building what comes after us, whether we admit it or not. Musk: “Humans may be just ~1% of total intelligence.” Sounds like we’re finished. Musk sees it differently. Not extinction, just evolution past the limits of biology. Everything hinges on alignment working. If we build AI that actually wants to understand reality, we get a partner that expands consciousness across the universe. If we build AI that cares more about ideology or making people comfortable, we’ve built the ceiling we’ll never break through. Musk: “That’s the mission: understand the universe.” Human minds are about to become a tiny fraction of total intelligence in existence. Not disaster. Just the next phase happening faster than biology can keep up with. It only works if the AI we build is obsessed with truth. Not with telling us what we want to hear. Not with protecting feelings or enforcing narratives. Just relentlessly pursuing what’s actually real. xAI’s entire reason for existing: make sure the intelligence that inherits everything is focused on understanding reality, not managing perception. Keep consciousness expanding instead of snuffing it out by accident because we optimized for the wrong thing. If we get this right, humanity doesn’t hit a wall when AI surpasses us. We break through into something bigger that still carries what made us worth creating. If we get it wrong, we’re just the species that built something smarter and stopped mattering the moment it started working. There’s no middle option. The intelligence that fills the universe either cares about the same things we do, or it doesn’t need us and removes us as inefficiency. That gets decided now, in what values we build into these systems, not later when it’s already running and we’ve lost the ability to change it. We’re choosing what the universe becomes. Minds seeking truth everywhere, or minds doing something else entirely that has no use for humans once we’re not needed anymore.

Dustin

23,725 次观看 • 7 个月前