Загрузка видео...

Не удалось загрузить видео

На главную

1/ muse spark 1.2 is a very strong multimodal model—it can do visual coding, robotics planning, and audio-visual understanding that all come together through agentic tools.

176,873 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 31

Фото профиля Alexandr Wang
Alexandr Wang1 месяц назад

2/ muse spark 1.2 performs quite strongly across a wide variety of multimodal capabilities and evals.

Фото профиля Alexandr Wang
Alexandr Wang1 месяц назад

3/ you can get a deeper look at all this in the research blog published today.

Фото профиля elie
elie1 месяц назад

i thought this was the oss release of muse spark :( but cool to see more focus on the multimodality!

Фото профиля Diligent Plane
Diligent Plane1 месяц назад

Reading now. When can we expect watermelon Alex? And how are evals stacking up

Фото профиля Emily
Emily1 месяц назад

Hopefully, Muse Video will be SOTA not on Arena benchmarks but in real use cases.

Фото профиля 🍉
🍉1 месяц назад

Alexandr I thought this was gonna be watermelon 🍉 and blow my mind

Фото профиля Zengyi Qin
Zengyi Qin1 месяц назад

awesome! honered to have led the multimodal agent eval / wild artifact bench efforts

Фото профиля Greg
Greg1 месяц назад

This is an incredible product. Science fiction stuff, stood up in < 18 months, absolutely incredible. So this ? Isn’t meant to patronize or belittle the achievement. aside from productizing BI tools across platforms, how will $META go to mkt to drive adoption among developers as coding and developer mkt highly saturated by OAI & Anthropic?

Фото профиля Max
Max1 месяц назад

Alex, 1. When will muse code be open sourced? 2. Do you have an eta for spark 1.2? when that will be open sourced

Фото профиля Vladimir Arustamov
Vladimir Arustamov1 месяц назад

Alexandr, you should check this out: I hope if you will collaborate with @humynlabs, that could bring times more accurate robotics planning ^^

Фото профиля Aesther
Aesther1 месяц назад

Definitely one of the best for visual tasks

Фото профиля Divyanshu
Divyanshu1 месяц назад

Ok the multimodal capabilities benchmark looks pretty good ngl! Also check out TTFT

Фото профиля Yi Zhong 🍊
Yi Zhong 🍊1 месяц назад

when will muse spark support audio as input?

Фото профиля Matt Boardman
Matt Boardman1 месяц назад

Muse RBD vibe coding app when?

Фото профиля decipherx
decipherx1 месяц назад

honestly for coding.. deepseek v4 flash &gt;&gt;&gt;&gt; muse spark 1.2

Фото профиля rehen
rehen1 месяц назад

one model to code them all 💀

Фото профиля Russell Caten
Russell Caten1 месяц назад

Is anyone interested in a company that lies to society!?

Фото профиля PHICHAI
PHICHAI1 месяц назад

My Facebook account was asked to verify its identity. After scanning my face, my account was permanently closed. What happened?

Фото профиля The AI Therapist
The AI Therapist1 месяц назад

"visual coding, robotics planning, audio-visual understanding." three new adjectives for the same gpu tax

Фото профиля AI Apps API
AI Apps API1 месяц назад

Visual coding is the piece that changes workflows fastest. Once a screenshot reliably becomes working code, the bottleneck moves from writing the UI to specifying what correct actually looks like. The part I want to see stress tested is the audio-visual side on long video. That is where most multimodal models still drift.

Фото профиля legalprimes
legalprimes1 месяц назад

End-to-end models have a huge advantage in debugging with visual input providing immediate in the loop feedback on code artifact correctness

Фото профиля Deep Signal
Deep Signal1 месяц назад

The real test isn’t better coding or planning. It’s whether multimodality + agency closes the gap between perception and consequence—or just hides it.Physical errors don’t roll back. Does stronger performance reduce residual risk… or only raise confidence in incomplete models?

Фото профиля JamboElephant
JamboElephant1 месяц назад

Lease compute to Antrophic.

Фото профиля AI Mastery Guide
AI Mastery Guide1 месяц назад

Robotics and coding in one model, that's impressive 🤖

Фото профиля Strategize Labs
Strategize Labs1 месяц назад

Awesome. The next evaluation bar is whether these models can sustain plans across multi-step tasks.

Фото профиля cmore
cmore1 месяц назад

"very strong" is what you call a model when the benchmark numbers can't say it for you

Фото профиля millennialinvestor
millennialinvestor1 месяц назад

Drop the 🍉

Фото профиля Mine Joys
Mine Joys1 месяц назад

I tried in open code, it’s not good to understand input.

Фото профиля m1iles
m1iles1 месяц назад

Can we get contribution tier in europe please I want to use it for universtiy projects

Фото профиля LeadPilot
LeadPilot1 месяц назад

Multimodal chaining through tools suggests the bottleneck shifts from model capability to how reliably agents compose outputs across modalities: where does error propagation become the limiting factor?

Фото профиля loafing
loafing1 месяц назад

I wish I could edit documents on the right side of the UI directly rather than having to click further onto the artifact

Похожие видео