Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Auditable by design. Every model call, tool run, and edit hits a local event log before it executes. If it crashes mid-task, it picks up exactly where it left off from that log. No lost work and no re-prompting.

91,554 görüntüleme • 1 ay önce •via X (Twitter)

19 Yorum

Mark Zuckerberg profil fotoğrafı
Mark Zuckerberg1 ay önce

We pointed Muse Spark 1.2 at a kernel optimization task and let it run. 1,000+ tool calls over 24 hours on NVIDIA Hopper. It kept finding substantial improvements well beyond the initial exploration phase.

Mark Zuckerberg profil fotoğrafı
Mark Zuckerberg1 ay önce

Muse Code runs specialized background agents that stay active your whole session, so they build up context over time instead of starting from scratch on every task. When a job is big enough, it fans out to separate sub-agents working in parallel in isolated worktrees. Your working copy is never touched. In testing we had it build six features for a game simultaneously with no collisions.

Mark Zuckerberg profil fotoğrafı
Mark Zuckerberg1 ay önce

Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.

Mark Zuckerberg profil fotoğrafı
Mark Zuckerberg1 ay önce

Pricing: It's easy and low-cost to get started. Install Muse Code with one line and you can start on our contributor tier.

Mark Zuckerberg profil fotoğrafı
Mark Zuckerberg1 ay önce

Muse Spark 1.2 is our next step as we push toward frontier, with larger, more capable models on the way. Install it, use it, tell us what you think.

Evil Rabbit profil fotoğrafı
Evil Rabbit1 ay önce

impressive Mark! I’ll have to give it a try

Sarjan Narwan profil fotoğrafı
Sarjan Narwan1 ay önce

I wonder if that'd cause issues with non indempotent operations.

Olli profil fotoğrafı
Olli1 ay önce

A local event log changes the shape of a coding agent. It makes the run inspectable as a system process, not only as a chat. Resume is useful, but I would judge it by whether a reviewer can trace each tool call, permission, file change and rollback after the task is done.

🇺🇸 Michael Childress 🇺🇸 profil fotoğrafı
🇺🇸 Michael Childress 🇺🇸1 ay önce

I Have Figured Out How To Power The Cooling Of New Data Center Constructions, I USE NO WATER, I Also Cover 10-15 Percent Of The Server Energy, Wanting To Make The Most Powerful Data Centers On Earth, Please Help Me Help You Become The Most Powerful AI Company On Earth, Very Serious, Please Just Listen To My Ideas Is All ...

rempred profil fotoğrafı
rempred1 ay önce

It's funny to see something I came up with over a month ago get implemented in a major frontier lab

Sarah Khan profil fotoğrafı
Sarah Khan1 ay önce

Hello @finkd @instagram We don't need new features. Fix your AI and bring back real human support. All my accounts have been suspended for over 2 months, Please help 💔

Marcin Wójcik 🎰 profil fotoğrafı
Marcin Wójcik 🎰1 ay önce

Lately, I’ve come to the conclusion that working for Mr. Zuckerberg on building metacognitive structures might be the most productive path. After all, it is the only company that doesn’t churn out as much narrative as the others. (If<What if)

Happy Tails profil fotoğrafı
Happy Tails1 ay önce

State persistence is the biggest bottleneck holding back true AI agents right now. Pick up where it left off without burning tokens on re-prompting? Game changer for long-running workflows.

M-viz profil fotoğrafı
M-viz1 ay önce

That's how serious AI systems should be built. An assistant that can recover from failures without losing context feels far more reliable than one that starts over every time. Auditability isn't just for compliance—it improves the user experience too.

Giedrius Trump profil fotoğrafı
Giedrius Trump1 ay önce

That’s nice

Tornado guy profil fotoğrafı
Tornado guy1 ay önce

Such a nice output

Aj Castelletto profil fotoğrafı
Aj Castelletto1 ay önce

Hey Mark, don't mind me, just shamelessly plugging my book preorder under what's sure to be a banger of a post. Did I mention the main characters name is Muse?!

Tim Kostolansky profil fotoğrafı
Tim Kostolansky1 ay önce

video unrelated

Jeff James Martin profil fotoğrafı
Jeff James Martin1 ay önce

Auditability is going to be one of the dividing lines between AI that demos well and AI that organizations can actually trust with real work. When leaders can see what happened, why it happened and where ownership sits, autonomy becomes much easier to scale.

Benzer Videolar

Harness vs. Graphs, clearly explained! a harness is great, and most people think it is the whole thing: retries, timeouts, a sandbox, a log, the context it assembles before every call. all of that is real work, and all of it wraps exactly one call. run it a hundred times and you have one call, made very safely, a hundred times. Graph engineering fixes this by moving the decision up a layer: not how safely one call is made, but which calls exist to be made at all. you need both, and here is the sentence that resolves the whole confusion: the harness is everything around one call. the graph is everything between them. ↳ around one call: retry, timeout, sandbox, log, assemble the context, hand back a result ↳ between calls: split, fan out, merge, gate, send back Prompts → Context → Harness → Loops → Graphs the harness does not go away when you build a graph. it moves under each node, and now there are five of them, each wrapping a call you would never have made by hand. the trick is knowing which layer a failure belongs to. turn a piece off and run it again. if the call still works, it was the harness. if the wrong step runs at all, it was the graph. people spend weeks hardening a harness around a node that should not have existed. one thing to know before you scale it. most of what people call their agent is a harness with a chat box on it. ↳ it retries, it times out, it logs, it assembles context, it holds one call up beautifully ↳ it has never once decided that a second call should exist, and that is the entire difference that last one catches careful people. a harness that never fails is not evidence the system is right. it is evidence one call went well, which is the smallest possible claim. and the one that eats whole nights: a harness cannot save you from the wrong step running. you can retry a bad decision three times with a clean log and perfect isolation, and all you bought was three copies of it. below i have quoted my full guide on graph engineering. it covers the three topologies, the verifier patterns, and where the gate should actually open. save this and read it below ↓

Hanako

51,536 görüntüleme • 20 gün önce