Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

1/ muse spark 1.2 is a very strong multimodal model—it can do visual coding, robotics planning, and audio-visual understanding that all come together through agentic tools.

176,873 Aufrufe • vor 1 Monat •via X (Twitter)

31 Kommentare

Profilbild von Alexandr Wang
Alexandr Wangvor 1 Monat

2/ muse spark 1.2 performs quite strongly across a wide variety of multimodal capabilities and evals.

Profilbild von Alexandr Wang
Alexandr Wangvor 1 Monat

3/ you can get a deeper look at all this in the research blog published today.

Profilbild von elie
elievor 1 Monat

i thought this was the oss release of muse spark :( but cool to see more focus on the multimodality!

Profilbild von Diligent Plane
Diligent Planevor 1 Monat

Reading now. When can we expect watermelon Alex? And how are evals stacking up

Profilbild von Emily
Emilyvor 1 Monat

Hopefully, Muse Video will be SOTA not on Arena benchmarks but in real use cases.

Profilbild von 🍉
🍉vor 1 Monat

Alexandr I thought this was gonna be watermelon 🍉 and blow my mind

Profilbild von Zengyi Qin
Zengyi Qinvor 1 Monat

awesome! honered to have led the multimodal agent eval / wild artifact bench efforts

Profilbild von Greg
Gregvor 1 Monat

This is an incredible product. Science fiction stuff, stood up in < 18 months, absolutely incredible. So this ? Isn’t meant to patronize or belittle the achievement. aside from productizing BI tools across platforms, how will $META go to mkt to drive adoption among developers as coding and developer mkt highly saturated by OAI & Anthropic?

Profilbild von Max
Maxvor 1 Monat

Alex, 1. When will muse code be open sourced? 2. Do you have an eta for spark 1.2? when that will be open sourced

Profilbild von Vladimir Arustamov
Vladimir Arustamovvor 1 Monat

Alexandr, you should check this out: I hope if you will collaborate with @humynlabs, that could bring times more accurate robotics planning ^^

Profilbild von Aesther
Aesthervor 1 Monat

Definitely one of the best for visual tasks

Profilbild von Divyanshu
Divyanshuvor 1 Monat

Ok the multimodal capabilities benchmark looks pretty good ngl! Also check out TTFT

Profilbild von Yi Zhong 🍊
Yi Zhong 🍊vor 1 Monat

when will muse spark support audio as input?

Profilbild von Matt Boardman
Matt Boardmanvor 1 Monat

Muse RBD vibe coding app when?

Profilbild von decipherx
decipherxvor 1 Monat

honestly for coding.. deepseek v4 flash &gt;&gt;&gt;&gt; muse spark 1.2

Profilbild von rehen
rehenvor 1 Monat

one model to code them all 💀

Profilbild von Russell Caten
Russell Catenvor 1 Monat

Is anyone interested in a company that lies to society!?

Profilbild von PHICHAI
PHICHAIvor 1 Monat

My Facebook account was asked to verify its identity. After scanning my face, my account was permanently closed. What happened?

Profilbild von The AI Therapist
The AI Therapistvor 1 Monat

"visual coding, robotics planning, audio-visual understanding." three new adjectives for the same gpu tax

Profilbild von AI Apps API
AI Apps APIvor 1 Monat

Visual coding is the piece that changes workflows fastest. Once a screenshot reliably becomes working code, the bottleneck moves from writing the UI to specifying what correct actually looks like. The part I want to see stress tested is the audio-visual side on long video. That is where most multimodal models still drift.

Profilbild von legalprimes
legalprimesvor 1 Monat

End-to-end models have a huge advantage in debugging with visual input providing immediate in the loop feedback on code artifact correctness

Profilbild von Deep Signal
Deep Signalvor 1 Monat

The real test isn’t better coding or planning. It’s whether multimodality + agency closes the gap between perception and consequence—or just hides it.Physical errors don’t roll back. Does stronger performance reduce residual risk… or only raise confidence in incomplete models?

Profilbild von JamboElephant
JamboElephantvor 1 Monat

Lease compute to Antrophic.

Profilbild von AI Mastery Guide
AI Mastery Guidevor 1 Monat

Robotics and coding in one model, that's impressive 🤖

Profilbild von Strategize Labs
Strategize Labsvor 1 Monat

Awesome. The next evaluation bar is whether these models can sustain plans across multi-step tasks.

Profilbild von cmore
cmorevor 1 Monat

"very strong" is what you call a model when the benchmark numbers can't say it for you

Profilbild von millennialinvestor
millennialinvestorvor 1 Monat

Drop the 🍉

Profilbild von Mine Joys
Mine Joysvor 1 Monat

I tried in open code, it’s not good to understand input.

Profilbild von m1iles
m1ilesvor 1 Monat

Can we get contribution tier in europe please I want to use it for universtiy projects

Profilbild von LeadPilot
LeadPilotvor 1 Monat

Multimodal chaining through tools suggests the bottleneck shifts from model capability to how reliably agents compose outputs across modalities: where does error propagation become the limiting factor?

Profilbild von loafing
loafingvor 1 Monat

I wish I could edit documents on the right side of the UI directly rather than having to click further onto the artifact

Ähnliche Videos