正在加载视频...

视频加载失败

1/ muse spark 1.2 is a very strong multimodal model—it can do visual coding, robotics planning, and audio-visual understanding that all come together through agentic tools.

176,873 次观看 • 1 个月前 •via X (Twitter)

31 条评论

Alexandr Wang 的头像
Alexandr Wang1 个月前

2/ muse spark 1.2 performs quite strongly across a wide variety of multimodal capabilities and evals.

Alexandr Wang 的头像
Alexandr Wang1 个月前

3/ you can get a deeper look at all this in the research blog published today.

elie 的头像
elie1 个月前

i thought this was the oss release of muse spark :( but cool to see more focus on the multimodality!

Diligent Plane 的头像
Diligent Plane1 个月前

Reading now. When can we expect watermelon Alex? And how are evals stacking up

Emily 的头像
Emily1 个月前

Hopefully, Muse Video will be SOTA not on Arena benchmarks but in real use cases.

🍉 的头像
🍉1 个月前

Alexandr I thought this was gonna be watermelon 🍉 and blow my mind

Zengyi Qin 的头像
Zengyi Qin1 个月前

awesome! honered to have led the multimodal agent eval / wild artifact bench efforts

Greg 的头像
Greg1 个月前

This is an incredible product. Science fiction stuff, stood up in < 18 months, absolutely incredible. So this ? Isn’t meant to patronize or belittle the achievement. aside from productizing BI tools across platforms, how will $META go to mkt to drive adoption among developers as coding and developer mkt highly saturated by OAI & Anthropic?

Max 的头像
Max1 个月前

Alex, 1. When will muse code be open sourced? 2. Do you have an eta for spark 1.2? when that will be open sourced

Vladimir Arustamov 的头像
Vladimir Arustamov1 个月前

Alexandr, you should check this out: I hope if you will collaborate with @humynlabs, that could bring times more accurate robotics planning ^^

Aesther 的头像
Aesther1 个月前

Definitely one of the best for visual tasks

Divyanshu 的头像
Divyanshu1 个月前

Ok the multimodal capabilities benchmark looks pretty good ngl! Also check out TTFT

Yi Zhong 🍊 的头像
Yi Zhong 🍊1 个月前

when will muse spark support audio as input?

Matt Boardman 的头像
Matt Boardman1 个月前

Muse RBD vibe coding app when?

decipherx 的头像
decipherx1 个月前

honestly for coding.. deepseek v4 flash &gt;&gt;&gt;&gt; muse spark 1.2

rehen 的头像
rehen1 个月前

one model to code them all 💀

Russell Caten 的头像
Russell Caten1 个月前

Is anyone interested in a company that lies to society!?

PHICHAI 的头像
PHICHAI1 个月前

My Facebook account was asked to verify its identity. After scanning my face, my account was permanently closed. What happened?

The AI Therapist 的头像
The AI Therapist1 个月前

"visual coding, robotics planning, audio-visual understanding." three new adjectives for the same gpu tax

AI Apps API 的头像
AI Apps API1 个月前

Visual coding is the piece that changes workflows fastest. Once a screenshot reliably becomes working code, the bottleneck moves from writing the UI to specifying what correct actually looks like. The part I want to see stress tested is the audio-visual side on long video. That is where most multimodal models still drift.

legalprimes 的头像
legalprimes1 个月前

End-to-end models have a huge advantage in debugging with visual input providing immediate in the loop feedback on code artifact correctness

Deep Signal 的头像
Deep Signal1 个月前

The real test isn’t better coding or planning. It’s whether multimodality + agency closes the gap between perception and consequence—or just hides it.Physical errors don’t roll back. Does stronger performance reduce residual risk… or only raise confidence in incomplete models?

JamboElephant 的头像
JamboElephant1 个月前

Lease compute to Antrophic.

AI Mastery Guide 的头像
AI Mastery Guide1 个月前

Robotics and coding in one model, that's impressive 🤖

Strategize Labs 的头像
Strategize Labs1 个月前

Awesome. The next evaluation bar is whether these models can sustain plans across multi-step tasks.

cmore 的头像
cmore1 个月前

"very strong" is what you call a model when the benchmark numbers can't say it for you

millennialinvestor 的头像
millennialinvestor1 个月前

Drop the 🍉

Mine Joys 的头像
Mine Joys1 个月前

I tried in open code, it’s not good to understand input.

m1iles 的头像
m1iles1 个月前

Can we get contribution tier in europe please I want to use it for universtiy projects

LeadPilot 的头像
LeadPilot1 个月前

Multimodal chaining through tools suggests the bottleneck shifts from model capability to how reliably agents compose outputs across modalities: where does error propagation become the limiting factor?

loafing 的头像
loafing1 个月前

I wish I could edit documents on the right side of the UI directly rather than having to click further onto the artifact

相关视频