Загрузка видео...

Не удалось загрузить видео

На главную

Muse Spark 1.2 supports a broad range of multimodal tasks, from turning visuals into working code to translating perception into physical action. It also brings robust audio-visual understanding to enable video-heavy workflows common in real-world enterprise use. Today, we’re sharing new evals and demos that illustrate the breadth of...

58,163 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 35

Фото профиля AI at Meta
AI at Meta1 месяц назад

Now let’s take a look at how Muse Spark 1.2 performs across visual reasoning, chart understanding, and knowledge-intensive tasks.

Фото профиля AI at Meta
AI at Meta1 месяц назад

Muse Spark 1.2 reasons better with tool use. The model inspects visual inputs more closely and incorporates what it finds into its reasoning.

Фото профиля AI at Meta
AI at Meta1 месяц назад

Muse Spark 1.2 generates digital artifacts like web pages and games directly from images or video. It translates visual layout, hierarchy, and style into working code, evaluating correctness based on actual rendering and behavior while using a continuous self-improvement loop to refine its outputs.

Фото профиля AI at Meta
AI at Meta1 месяц назад

Muse Spark brings spatial intelligence to robotics. A specialized variant of Muse Spark acts as the robot brain and orchestrator: it takes a user instruction, decodes tool calls, observes the results, and loops until the task is complete. As shown in this demo, Muse Spark 1.2 can plan sub-tasks for a bimanual robot tidying a desk, distinguishing a hair brush from a makeup brush and placing the lipstick in a drawer.

Фото профиля AI at Meta
AI at Meta1 месяц назад

Internally, Muse Spark is deployed in several areas including media generation. For example, the model works with Muse Image for agentic media generation and produces detailed captions as training data for Muse Image and Muse Video. Muse Spark is also able to translate raw text, image, and video content into signals and insights that can be used in downstream applications. See more Muse Spark 1.2 evals and demos here:

Фото профиля Atoof
Atoof1 месяц назад

@xiaolonw Teach my guy some CAD

Фото профиля Velvetere
Velvetere1 месяц назад

Can’t wait to replace my dumb employees who can’t speak English with one of these bad boys. Keep it up

Фото профиля Apollo
Apollo1 месяц назад

Meta'sanswer to google's Gemini Robotics ER 2 ? it also processes images and video as input.

Фото профиля Clausius
Clausius1 месяц назад

@cursor_ai please support muse spark 1.2!

Фото профиля Default Settlement
Default Settlement1 месяц назад

When perception becomes physical action, a successful tool call is not proof of task completion. For this run, verification should bind the relevant observations, model version, authorized objective, navigation commands and final sensor evidence showing that the robot found the correct object. Multimodal capability expands what agents can perceive and do. Evidence rails prove what actually happened in the physical world.

Фото профиля Victoria Neiman
Victoria Neiman1 месяц назад

same model powering system-wide dictation on mac is now navigating robots through unstructured environments. muse spark isn't a chatbot – it's becoming the perception layer meta embeds into everything. text fields, ad dashboards, physical space

Фото профиля Gail Castillo
Gail Castillo1 месяц назад

Look stupid. Mark, you don't get it. In this AI competition, all you had to do to win was do nothing, like Apple. Just like nobody would take Google Hangouts seriously for work calls (or Zoom would never have had a chance), nobody would use Facebook products for real work.

Фото профиля Inflectiv AI ⧉
Inflectiv AI ⧉1 месяц назад

Translating raw multimodal perception directly into physical robotics control and spatial action is a huge step forward for real-world embodiment.

Фото профиля Fajar M Reza
Fajar M Reza1 месяц назад

Multimodal models translating visuals into code could shorten enterprise automation cycles.

Фото профиля Jack Winterland
Jack Winterland1 месяц назад

Can someone tell me what’s the closest estimated parameter count for Muse Spark 1.1?

Фото профиля Love Web3 World
Love Web3 World1 месяц назад

It’s AI turning perception into action with tools, code, and robots in the loop. Muse Spark 1.2 reportedly jumps from 59.8 to 72.0 with tool use.

Фото профиля seastart
seastart1 месяц назад

The jump from “understanding what’s in the video” to actually taking the next action is the interesting part. That’s when multimodal models start becoming useful infrastructure, not just better perception.

Фото профиля Ishraq Syed
Ishraq Syed1 месяц назад

Can't wait to see muse video.. 😍

Фото профиля Si(super intelligence)
Si(super intelligence)1 месяц назад

👍

Фото профиля Adhiraj Budhathoki
Adhiraj Budhathoki1 месяц назад

Consider this my final warning: I have documented every single failure with screen recordings. If my access isn't fixed immediately, I will be posting these recordings publicly for the world to see how broken your platform is. Fix this now. @MetaNewsroom

Фото профиля Chris.base
Chris.base1 месяц назад

curious to see how well it handles long, messy enterprise video beyond polished demos

Фото профиля 404 Opinions
404 Opinions1 месяц назад

robot finds duck meta finds way to make us pay for it

Фото профиля tavarke
tavarke29 дней назад

Meta has become unreliable

Фото профиля Eva Morgan
Eva Morgan1 месяц назад

The combination of vision, reasoning, and tool use opens up exciting possibilities for robotics and enterprise automation.

Фото профиля Kostya | AI
Kostya | AI1 месяц назад

Muse Spark 1.2’s ability to turn visual inputs directly into executable code could cut prototype cycles by orders of magnitude for data‑science teams.

Фото профиля Asli Al-Koul
Asli Al-Koul1 месяц назад

ممكن ادعم من كلشي

Фото профиля メEscanorメ
メEscanorメ1 месяц назад

The jump from understanding visuals to taking real-world action is where multimodal AI gets truly interesting.

Фото профиля 아논
아논1 месяц назад

한국에서도 쓸 수 있는 날이 오나

Фото профиля 美股搬运工(诚信交朋友中)
美股搬运工(诚信交朋友中)1 месяц назад

can muse ouput arm actions to complete task as folds?

Фото профиля Sharjeel Younas
Sharjeel Younas1 месяц назад

This is where multimodal AI gets really powerful—connecting vision, reasoning, and tool use to drive physical actions in real-world environments.

Фото профиля Asli Al-Koul
Asli Al-Koul1 месяц назад

ممكن ادعم

Фото профиля SimpleMan
SimpleMan26 дней назад

Enable Instagram Account quizmasterdhan

Фото профиля Sophie Walker
Sophie Walker1 месяц назад

Curious if the audio-visual understanding truly handles messy real-world video or just clean demos. The evals sound promising, but I want proof!

Фото профиля Ricci Research
Ricci Research1 месяц назад

"Sequences have been shortened throughout" is the most informative caption on the screen — the gap between demo-time and wall-clock time is the whole story in embodied agents right now.

Фото профиля Alex Zverianskii
Alex Zverianskii1 месяц назад

Loop until the task is complete is the contract. Progress as state, not first token.

Похожие видео