Загрузка видео...

Не удалось загрузить видео

На главную

Auditable by design. Every model call, tool run, and edit hits a local event log before it executes. If it crashes mid-task, it picks up exactly where it left off from that log. No lost work and no re-prompting.

91,554 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 19

Фото профиля Mark Zuckerberg
Mark Zuckerberg1 месяц назад

We pointed Muse Spark 1.2 at a kernel optimization task and let it run. 1,000+ tool calls over 24 hours on NVIDIA Hopper. It kept finding substantial improvements well beyond the initial exploration phase.

Фото профиля Mark Zuckerberg
Mark Zuckerberg1 месяц назад

Muse Code runs specialized background agents that stay active your whole session, so they build up context over time instead of starting from scratch on every task. When a job is big enough, it fans out to separate sub-agents working in parallel in isolated worktrees. Your working copy is never touched. In testing we had it build six features for a game simultaneously with no collisions.

Фото профиля Mark Zuckerberg
Mark Zuckerberg1 месяц назад

Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.

Фото профиля Mark Zuckerberg
Mark Zuckerberg1 месяц назад

Pricing: It's easy and low-cost to get started. Install Muse Code with one line and you can start on our contributor tier.

Фото профиля Mark Zuckerberg
Mark Zuckerberg1 месяц назад

Muse Spark 1.2 is our next step as we push toward frontier, with larger, more capable models on the way. Install it, use it, tell us what you think.

Фото профиля Evil Rabbit
Evil Rabbit1 месяц назад

impressive Mark! I’ll have to give it a try

Фото профиля Sarjan Narwan
Sarjan Narwan1 месяц назад

I wonder if that'd cause issues with non indempotent operations.

Фото профиля Olli
Olli1 месяц назад

A local event log changes the shape of a coding agent. It makes the run inspectable as a system process, not only as a chat. Resume is useful, but I would judge it by whether a reviewer can trace each tool call, permission, file change and rollback after the task is done.

Фото профиля 🇺🇸 Michael Childress 🇺🇸
🇺🇸 Michael Childress 🇺🇸1 месяц назад

I Have Figured Out How To Power The Cooling Of New Data Center Constructions, I USE NO WATER, I Also Cover 10-15 Percent Of The Server Energy, Wanting To Make The Most Powerful Data Centers On Earth, Please Help Me Help You Become The Most Powerful AI Company On Earth, Very Serious, Please Just Listen To My Ideas Is All ...

Фото профиля rempred
rempred1 месяц назад

It's funny to see something I came up with over a month ago get implemented in a major frontier lab

Фото профиля Sarah Khan
Sarah Khan1 месяц назад

Hello @finkd @instagram We don't need new features. Fix your AI and bring back real human support. All my accounts have been suspended for over 2 months, Please help 💔

Фото профиля Marcin Wójcik 🎰
Marcin Wójcik 🎰1 месяц назад

Lately, I’ve come to the conclusion that working for Mr. Zuckerberg on building metacognitive structures might be the most productive path. After all, it is the only company that doesn’t churn out as much narrative as the others. (If<What if)

Фото профиля Happy Tails
Happy Tails1 месяц назад

State persistence is the biggest bottleneck holding back true AI agents right now. Pick up where it left off without burning tokens on re-prompting? Game changer for long-running workflows.

Фото профиля M-viz
M-viz1 месяц назад

That's how serious AI systems should be built. An assistant that can recover from failures without losing context feels far more reliable than one that starts over every time. Auditability isn't just for compliance—it improves the user experience too.

Фото профиля Giedrius Trump
Giedrius Trump1 месяц назад

That’s nice

Фото профиля Tornado guy
Tornado guy1 месяц назад

Such a nice output

Фото профиля Aj Castelletto
Aj Castelletto1 месяц назад

Hey Mark, don't mind me, just shamelessly plugging my book preorder under what's sure to be a banger of a post. Did I mention the main characters name is Muse?!

Фото профиля Tim Kostolansky
Tim Kostolansky1 месяц назад

video unrelated

Фото профиля Jeff James Martin
Jeff James Martin1 месяц назад

Auditability is going to be one of the dividing lines between AI that demos well and AI that organizations can actually trust with real work. When leaders can see what happened, why it happened and where ownership sits, autonomy becomes much easier to scale.

Похожие видео

Harness vs. Graphs, clearly explained! a harness is great, and most people think it is the whole thing: retries, timeouts, a sandbox, a log, the context it assembles before every call. all of that is real work, and all of it wraps exactly one call. run it a hundred times and you have one call, made very safely, a hundred times. Graph engineering fixes this by moving the decision up a layer: not how safely one call is made, but which calls exist to be made at all. you need both, and here is the sentence that resolves the whole confusion: the harness is everything around one call. the graph is everything between them. ↳ around one call: retry, timeout, sandbox, log, assemble the context, hand back a result ↳ between calls: split, fan out, merge, gate, send back Prompts → Context → Harness → Loops → Graphs the harness does not go away when you build a graph. it moves under each node, and now there are five of them, each wrapping a call you would never have made by hand. the trick is knowing which layer a failure belongs to. turn a piece off and run it again. if the call still works, it was the harness. if the wrong step runs at all, it was the graph. people spend weeks hardening a harness around a node that should not have existed. one thing to know before you scale it. most of what people call their agent is a harness with a chat box on it. ↳ it retries, it times out, it logs, it assembles context, it holds one call up beautifully ↳ it has never once decided that a second call should exist, and that is the entire difference that last one catches careful people. a harness that never fails is not evidence the system is right. it is evidence one call went well, which is the smallest possible claim. and the one that eats whole nights: a harness cannot save you from the wrong step running. you can retry a bad decision three times with a clean log and perfect isolation, and all you bought was three copies of it. below i have quoted my full guide on graph engineering. it covers the three topologies, the verifier patterns, and where the gate should actually open. save this and read it below ↓

Hanako

51,536 просмотров • 20 дней назад