Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Muse Spark 1.2 supports a broad range of multimodal tasks, from turning visuals into working code to translating perception into physical action. It also brings robust audio-visual understanding to enable video-heavy workflows common in real-world enterprise use. Today, we’re sharing new evals and demos that illustrate the breadth of...

58,163 Aufrufe • vor 1 Monat •via X (Twitter)

35 Kommentare

Profilbild von AI at Meta
AI at Metavor 1 Monat

Now let’s take a look at how Muse Spark 1.2 performs across visual reasoning, chart understanding, and knowledge-intensive tasks.

Profilbild von AI at Meta
AI at Metavor 1 Monat

Muse Spark 1.2 reasons better with tool use. The model inspects visual inputs more closely and incorporates what it finds into its reasoning.

Profilbild von AI at Meta
AI at Metavor 1 Monat

Muse Spark 1.2 generates digital artifacts like web pages and games directly from images or video. It translates visual layout, hierarchy, and style into working code, evaluating correctness based on actual rendering and behavior while using a continuous self-improvement loop to refine its outputs.

Profilbild von AI at Meta
AI at Metavor 1 Monat

Muse Spark brings spatial intelligence to robotics. A specialized variant of Muse Spark acts as the robot brain and orchestrator: it takes a user instruction, decodes tool calls, observes the results, and loops until the task is complete. As shown in this demo, Muse Spark 1.2 can plan sub-tasks for a bimanual robot tidying a desk, distinguishing a hair brush from a makeup brush and placing the lipstick in a drawer.

Profilbild von AI at Meta
AI at Metavor 1 Monat

Internally, Muse Spark is deployed in several areas including media generation. For example, the model works with Muse Image for agentic media generation and produces detailed captions as training data for Muse Image and Muse Video. Muse Spark is also able to translate raw text, image, and video content into signals and insights that can be used in downstream applications. See more Muse Spark 1.2 evals and demos here:

Profilbild von Atoof
Atoofvor 1 Monat

@xiaolonw Teach my guy some CAD

Profilbild von Velvetere
Velveterevor 1 Monat

Can’t wait to replace my dumb employees who can’t speak English with one of these bad boys. Keep it up

Profilbild von Apollo
Apollovor 1 Monat

Meta'sanswer to google's Gemini Robotics ER 2 ? it also processes images and video as input.

Profilbild von Clausius
Clausiusvor 1 Monat

@cursor_ai please support muse spark 1.2!

Profilbild von Default Settlement
Default Settlementvor 1 Monat

When perception becomes physical action, a successful tool call is not proof of task completion. For this run, verification should bind the relevant observations, model version, authorized objective, navigation commands and final sensor evidence showing that the robot found the correct object. Multimodal capability expands what agents can perceive and do. Evidence rails prove what actually happened in the physical world.

Profilbild von Victoria Neiman
Victoria Neimanvor 1 Monat

same model powering system-wide dictation on mac is now navigating robots through unstructured environments. muse spark isn't a chatbot – it's becoming the perception layer meta embeds into everything. text fields, ad dashboards, physical space

Profilbild von Gail Castillo
Gail Castillovor 1 Monat

Look stupid. Mark, you don't get it. In this AI competition, all you had to do to win was do nothing, like Apple. Just like nobody would take Google Hangouts seriously for work calls (or Zoom would never have had a chance), nobody would use Facebook products for real work.

Profilbild von Inflectiv AI ⧉
Inflectiv AI ⧉vor 1 Monat

Translating raw multimodal perception directly into physical robotics control and spatial action is a huge step forward for real-world embodiment.

Profilbild von Fajar M Reza
Fajar M Rezavor 1 Monat

Multimodal models translating visuals into code could shorten enterprise automation cycles.

Profilbild von Jack Winterland
Jack Winterlandvor 1 Monat

Can someone tell me what’s the closest estimated parameter count for Muse Spark 1.1?

Profilbild von Love Web3 World
Love Web3 Worldvor 1 Monat

It’s AI turning perception into action with tools, code, and robots in the loop. Muse Spark 1.2 reportedly jumps from 59.8 to 72.0 with tool use.

Profilbild von seastart
seastartvor 1 Monat

The jump from “understanding what’s in the video” to actually taking the next action is the interesting part. That’s when multimodal models start becoming useful infrastructure, not just better perception.

Profilbild von Ishraq Syed
Ishraq Syedvor 1 Monat

Can't wait to see muse video.. 😍

Profilbild von Si(super intelligence)
Si(super intelligence)vor 1 Monat

👍

Profilbild von Adhiraj Budhathoki
Adhiraj Budhathokivor 1 Monat

Consider this my final warning: I have documented every single failure with screen recordings. If my access isn't fixed immediately, I will be posting these recordings publicly for the world to see how broken your platform is. Fix this now. @MetaNewsroom

Profilbild von Chris.base
Chris.basevor 1 Monat

curious to see how well it handles long, messy enterprise video beyond polished demos

Profilbild von 404 Opinions
404 Opinionsvor 1 Monat

robot finds duck meta finds way to make us pay for it

Profilbild von tavarke
tavarkevor 29 Tagen

Meta has become unreliable

Profilbild von Eva Morgan
Eva Morganvor 1 Monat

The combination of vision, reasoning, and tool use opens up exciting possibilities for robotics and enterprise automation.

Profilbild von Kostya | AI
Kostya | AIvor 1 Monat

Muse Spark 1.2’s ability to turn visual inputs directly into executable code could cut prototype cycles by orders of magnitude for data‑science teams.

Profilbild von Asli Al-Koul
Asli Al-Koulvor 1 Monat

ممكن ادعم من كلشي

Profilbild von メEscanorメ
メEscanorメvor 1 Monat

The jump from understanding visuals to taking real-world action is where multimodal AI gets truly interesting.

Profilbild von 아논
아논vor 1 Monat

한국에서도 쓸 수 있는 날이 오나

Profilbild von 美股搬运工(诚信交朋友中)
美股搬运工(诚信交朋友中)vor 1 Monat

can muse ouput arm actions to complete task as folds?

Profilbild von Sharjeel Younas
Sharjeel Younasvor 1 Monat

This is where multimodal AI gets really powerful—connecting vision, reasoning, and tool use to drive physical actions in real-world environments.

Profilbild von Asli Al-Koul
Asli Al-Koulvor 1 Monat

ممكن ادعم

Profilbild von SimpleMan
SimpleManvor 26 Tagen

Enable Instagram Account quizmasterdhan

Profilbild von Sophie Walker
Sophie Walkervor 1 Monat

Curious if the audio-visual understanding truly handles messy real-world video or just clean demos. The evals sound promising, but I want proof!

Profilbild von Ricci Research
Ricci Researchvor 1 Monat

"Sequences have been shortened throughout" is the most informative caption on the screen — the gap between demo-time and wall-clock time is the whole story in embodied agents right now.

Profilbild von Alex Zverianskii
Alex Zverianskiivor 1 Monat

Loop until the task is complete is the contract. Progress as state, not first token.

Ähnliche Videos