Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

1/ muse spark 1.2 is a very strong multimodal model—it can do visual coding, robotics planning, and audio-visual understanding that all come together through agentic tools.

176,873 görüntüleme • 1 ay önce •via X (Twitter)

31 Yorum

Alexandr Wang profil fotoğrafı
Alexandr Wang1 ay önce

2/ muse spark 1.2 performs quite strongly across a wide variety of multimodal capabilities and evals.

Alexandr Wang profil fotoğrafı
Alexandr Wang1 ay önce

3/ you can get a deeper look at all this in the research blog published today.

elie profil fotoğrafı
elie1 ay önce

i thought this was the oss release of muse spark :( but cool to see more focus on the multimodality!

Diligent Plane profil fotoğrafı
Diligent Plane1 ay önce

Reading now. When can we expect watermelon Alex? And how are evals stacking up

Emily profil fotoğrafı
Emily1 ay önce

Hopefully, Muse Video will be SOTA not on Arena benchmarks but in real use cases.

🍉 profil fotoğrafı
🍉1 ay önce

Alexandr I thought this was gonna be watermelon 🍉 and blow my mind

Zengyi Qin profil fotoğrafı
Zengyi Qin1 ay önce

awesome! honered to have led the multimodal agent eval / wild artifact bench efforts

Greg profil fotoğrafı
Greg1 ay önce

This is an incredible product. Science fiction stuff, stood up in < 18 months, absolutely incredible. So this ? Isn’t meant to patronize or belittle the achievement. aside from productizing BI tools across platforms, how will $META go to mkt to drive adoption among developers as coding and developer mkt highly saturated by OAI & Anthropic?

Max profil fotoğrafı
Max1 ay önce

Alex, 1. When will muse code be open sourced? 2. Do you have an eta for spark 1.2? when that will be open sourced

Vladimir Arustamov profil fotoğrafı
Vladimir Arustamov1 ay önce

Alexandr, you should check this out: I hope if you will collaborate with @humynlabs, that could bring times more accurate robotics planning ^^

Aesther profil fotoğrafı
Aesther1 ay önce

Definitely one of the best for visual tasks

Divyanshu profil fotoğrafı
Divyanshu1 ay önce

Ok the multimodal capabilities benchmark looks pretty good ngl! Also check out TTFT

Yi Zhong 🍊 profil fotoğrafı
Yi Zhong 🍊1 ay önce

when will muse spark support audio as input?

Matt Boardman profil fotoğrafı
Matt Boardman1 ay önce

Muse RBD vibe coding app when?

decipherx profil fotoğrafı
decipherx1 ay önce

honestly for coding.. deepseek v4 flash &gt;&gt;&gt;&gt; muse spark 1.2

rehen profil fotoğrafı
rehen1 ay önce

one model to code them all 💀

Russell Caten profil fotoğrafı
Russell Caten1 ay önce

Is anyone interested in a company that lies to society!?

PHICHAI profil fotoğrafı
PHICHAI1 ay önce

My Facebook account was asked to verify its identity. After scanning my face, my account was permanently closed. What happened?

The AI Therapist profil fotoğrafı
The AI Therapist1 ay önce

"visual coding, robotics planning, audio-visual understanding." three new adjectives for the same gpu tax

AI Apps API profil fotoğrafı
AI Apps API1 ay önce

Visual coding is the piece that changes workflows fastest. Once a screenshot reliably becomes working code, the bottleneck moves from writing the UI to specifying what correct actually looks like. The part I want to see stress tested is the audio-visual side on long video. That is where most multimodal models still drift.

legalprimes profil fotoğrafı
legalprimes1 ay önce

End-to-end models have a huge advantage in debugging with visual input providing immediate in the loop feedback on code artifact correctness

Deep Signal profil fotoğrafı
Deep Signal1 ay önce

The real test isn’t better coding or planning. It’s whether multimodality + agency closes the gap between perception and consequence—or just hides it.Physical errors don’t roll back. Does stronger performance reduce residual risk… or only raise confidence in incomplete models?

JamboElephant profil fotoğrafı
JamboElephant1 ay önce

Lease compute to Antrophic.

AI Mastery Guide profil fotoğrafı
AI Mastery Guide1 ay önce

Robotics and coding in one model, that's impressive 🤖

Strategize Labs profil fotoğrafı
Strategize Labs1 ay önce

Awesome. The next evaluation bar is whether these models can sustain plans across multi-step tasks.

cmore profil fotoğrafı
cmore1 ay önce

"very strong" is what you call a model when the benchmark numbers can't say it for you

millennialinvestor profil fotoğrafı
millennialinvestor1 ay önce

Drop the 🍉

Mine Joys profil fotoğrafı
Mine Joys1 ay önce

I tried in open code, it’s not good to understand input.

m1iles profil fotoğrafı
m1iles1 ay önce

Can we get contribution tier in europe please I want to use it for universtiy projects

LeadPilot profil fotoğrafı
LeadPilot1 ay önce

Multimodal chaining through tools suggests the bottleneck shifts from model capability to how reliably agents compose outputs across modalities: where does error propagation become the limiting factor?

loafing profil fotoğrafı
loafing1 ay önce

I wish I could edit documents on the right side of the UI directly rather than having to click further onto the artifact

Benzer Videolar