Загрузка видео...
Не удалось загрузить видео
Muse Spark 1.2 supports a broad range of multimodal tasks, from turning visuals into working code to translating perception into physical action. It also brings robust audio-visual understanding to enable video-heavy workflows common in real-world enterprise use. Today, we’re sharing new evals and demos that illustrate the breadth of... show more
58,163 просмотров • 1 месяц назад •via X (Twitter)
Комментарии: 35

Now let’s take a look at how Muse Spark 1.2 performs across visual reasoning, chart understanding, and knowledge-intensive tasks.

Muse Spark 1.2 reasons better with tool use. The model inspects visual inputs more closely and incorporates what it finds into its reasoning.

Muse Spark 1.2 generates digital artifacts like web pages and games directly from images or video. It translates visual layout, hierarchy, and style into working code, evaluating correctness based on actual rendering and behavior while using a continuous self-improvement loop to refine its outputs.

Muse Spark brings spatial intelligence to robotics. A specialized variant of Muse Spark acts as the robot brain and orchestrator: it takes a user instruction, decodes tool calls, observes the results, and loops until the task is complete. As shown in this demo, Muse Spark 1.2 can plan sub-tasks for a bimanual robot tidying a desk, distinguishing a hair brush from a makeup brush and placing the lipstick in a drawer.

Internally, Muse Spark is deployed in several areas including media generation. For example, the model works with Muse Image for agentic media generation and produces detailed captions as training data for Muse Image and Muse Video. Muse Spark is also able to translate raw text, image, and video content into signals and insights that can be used in downstream applications. See more Muse Spark 1.2 evals and demos here:

@xiaolonw Teach my guy some CAD

Can’t wait to replace my dumb employees who can’t speak English with one of these bad boys. Keep it up

Meta'sanswer to google's Gemini Robotics ER 2 ? it also processes images and video as input.

@cursor_ai please support muse spark 1.2!

When perception becomes physical action, a successful tool call is not proof of task completion. For this run, verification should bind the relevant observations, model version, authorized objective, navigation commands and final sensor evidence showing that the robot found the correct object. Multimodal capability expands what agents can perceive and do. Evidence rails prove what actually happened in the physical world.

same model powering system-wide dictation on mac is now navigating robots through unstructured environments. muse spark isn't a chatbot – it's becoming the perception layer meta embeds into everything. text fields, ad dashboards, physical space

Look stupid. Mark, you don't get it. In this AI competition, all you had to do to win was do nothing, like Apple. Just like nobody would take Google Hangouts seriously for work calls (or Zoom would never have had a chance), nobody would use Facebook products for real work.

Translating raw multimodal perception directly into physical robotics control and spatial action is a huge step forward for real-world embodiment.

Multimodal models translating visuals into code could shorten enterprise automation cycles.

Can someone tell me what’s the closest estimated parameter count for Muse Spark 1.1?

It’s AI turning perception into action with tools, code, and robots in the loop. Muse Spark 1.2 reportedly jumps from 59.8 to 72.0 with tool use.

The jump from “understanding what’s in the video” to actually taking the next action is the interesting part. That’s when multimodal models start becoming useful infrastructure, not just better perception.

Can't wait to see muse video.. 😍

👍

Consider this my final warning: I have documented every single failure with screen recordings. If my access isn't fixed immediately, I will be posting these recordings publicly for the world to see how broken your platform is. Fix this now. @MetaNewsroom

curious to see how well it handles long, messy enterprise video beyond polished demos

robot finds duck meta finds way to make us pay for it

Meta has become unreliable

The combination of vision, reasoning, and tool use opens up exciting possibilities for robotics and enterprise automation.

Muse Spark 1.2’s ability to turn visual inputs directly into executable code could cut prototype cycles by orders of magnitude for data‑science teams.

ممكن ادعم من كلشي

The jump from understanding visuals to taking real-world action is where multimodal AI gets truly interesting.

한국에서도 쓸 수 있는 날이 오나

can muse ouput arm actions to complete task as folds?

This is where multimodal AI gets really powerful—connecting vision, reasoning, and tool use to drive physical actions in real-world environments.

ممكن ادعم

Enable Instagram Account quizmasterdhan

Curious if the audio-visual understanding truly handles messy real-world video or just clean demos. The evals sound promising, but I want proof!

"Sequences have been shortened throughout" is the most informative caption on the screen — the gap between demo-time and wall-clock time is the whole story in embodied agents right now.

Loop until the task is complete is the contract. Progress as state, not first token.


