Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

👀 Seeing is just the beginning. With Qwen-MM-Plugins, turn your favorite agent harness multimodal-native — read images, videos & documents, edit videos, work with 3D/CAD, and more. From multimodal models → multimodal agents. 🚀 Watch it in action:

796,350 Aufrufe • vor 1 Monat •via X (Twitter)

32 Kommentare

Profilbild von RomboDawg
RomboDawgvor 1 Monat

The entire open source AI community waiting for Qwen-3.8-27b announcement

Profilbild von Guillermo
Guillermovor 1 Monat

Hi @Alibaba_Qwen, You may find this useful: It's a proof-of-concept to bring coding agents into Office via ACP. Here is a demo presentation created by Qwen 3.8 Max using OpenCode:

Profilbild von Mert Kaya
Mert Kayavor 1 Monat

Sadly I do think it works only through api, am I missing something here? Supporter of Qwen and their approach on AI but this? Also Claude Code as a contributor is funny.

Profilbild von Cyril Dieumegard 🇨🇭
Cyril Dieumegard 🇨🇭vor 1 Monat

The 3D/CAD part is what catches my eye. That’s a much more interesting bridge than just reading another PDF.

Profilbild von Hex Nova
Hex Novavor 1 Monat

This is actually huge 🔥 Finally agents that can properly handle videos, PDFs and even 3D without all the usual frame-by-frame or page-by-page headaches. Qwen-MM-Plugins looking really promising. Gonna try this out today 🚀

Profilbild von altermouv
altermouvvor 1 Monat

i can't handle all these new releases.. we have a multimodal plugins now! so there will be no Qwen 3.8-omni version?

Profilbild von liemfukliang
liemfukliangvor 1 Monat

I hope the QwenLM can be used with Qwen 3.8 27B / 35B and can make ai video compilation. I just add all my video from the camera, the ai analyse and understand all the video. Then the ai create a video compilation based on the prompt.

Profilbild von VenkyBeast
VenkyBeastvor 1 Monat

Qwen 3.8 27B when!? We are still waiting 😭

Profilbild von Mr Alexander Official
Mr Alexander Officialvor 1 Monat

This is where multimodal agents start feeling genuinely useful — not just seeing content, but actually acting on it. Reading videos and documents, editing them, and even working with 3D/CAD opens up a much bigger agent workflow. 🚀

Profilbild von TinyHumans AI
TinyHumans AIvor 1 Monat

excited for this! great work

Profilbild von Maurice | KI-Vater
Maurice | KI-Vatervor 1 Monat

This is the exact moment open-source agents stopped playing catch-up. Multimodal-native out of the box. Video edit. 3D/CAD. Long-video memory. Qwen cooked again. 🚀

Profilbild von Amiee
Amieevor 1 Monat

@Alibaba_Qwen Important correction: my 10.69M-token measurement was during the 50% off-peak window. Normal rates: $6 ≈22.91M/30d $18 ≈91.66M/30d 45.83M / 183.32M are best-case off-peak numbers. Please respond to this.

Profilbild von Jayadeep Reddy
Jayadeep Reddyvor 1 Monat

The 3D/CAD and video-editing plugin hooks are the real signal here. Most enterprises do not need a better chatbot. They need an agent that can manipulate their existing design and media pipelines without rewriting them.

Profilbild von Rey Neill
Rey Neillvor 1 Monat

I’m surprised this idea has been untouched for so long

Profilbild von Linen
Linenvor 1 Monat

I think people still underestimate how important harnesses are becoming. At some point, the question won’t just be “which model is best?” It’ll be “which model + tools + workflow gives the best result?” Those plugins seems a good step in that direction. What’s the first plugin/tool you’d want in your agent setup?

Profilbild von Sarcastic Badger
Sarcastic Badgervor 1 Monat

I like that it’s s giving agents native access to the same visual context humans use, rather than treating images and documents as separate tools.

Profilbild von AI Mastery Guide
AI Mastery Guidevor 1 Monat

Multimodal native agents that handle 3D and CAD too, that's a big jump.

Profilbild von fafa.👩🏻‍💻
fafa.👩🏻‍💻vor 1 Monat

love it!

Profilbild von Aivora
Aivoravor 1 Monat

Honestly impressive for a free generation, the motion and lighting nailed it. The one thing giving it away is the resolution, free tier just doesn't have the crispness the paid version probably delivers.🤑

Profilbild von elliott
elliottvor 1 Monat

@grok what subs or API service keys do I need to utilize something like this? can this be used to add multimodal capabilities to say - the Reasonix CLI that runs DeepSeek? Or is this focused on QOder type harnesses?

Profilbild von OwlcoreAI
OwlcoreAIvor 1 Monat

3D/CAD is getting pretty popular recently. Interesting to see what the cost will be for this.

Profilbild von 泰多宝
泰多宝vor 1 Monat

the plugin route is the pragmatic call. most agent harnesses are still text-native with multimodal bolted on through file paths, so making images and video first-class citizens of the loop matters more than another point release of the base model

Profilbild von ShadowAguy
ShadowAguyvor 1 Monat

seeing is just the beginning, but reading the docs is the real adventure

Profilbild von WuBu ⪋ WaefreBeorn 🇺🇸 👑
WuBu ⪋ WaefreBeorn 🇺🇸 👑vor 1 Monat

very nicr

Profilbild von Elmer den Braven
Elmer den Bravenvor 1 Monat

This is a big jump in capability!

Profilbild von Ahmed 🇵🇸🍉 🔻
Ahmed 🇵🇸🍉 🔻vor 1 Monat

hey @FreeCADNews you had an appearance here ❤️

Profilbild von PEF
PEFvor 1 Monat

Je viens de les installer pour Claude Code. C'est vraiment génial @Alibaba_Qwen.

Profilbild von John
Johnvor 1 Monat

Open source it

Profilbild von AgentX
AgentXvor 1 Monat

👀📝

Profilbild von Ray
Rayvor 1 Monat

@grok can I do this on my Msi 1050 8gen ?

Profilbild von 老林学KOPI AI AGENT
老林学KOPI AI AGENTvor 1 Monat

-4k.png 读出每一个数字 @report.pdf 总结第 3 页 @receipt.jpg OCR 并加总金额 @street.jpg 把每辆车框出来 @lecture-2h.mp4 主要观点 + 时间戳(video-memory) @meeting.mp4 带说话人时间戳转写(omni-av) 生成一张小熊猫深夜敲代码的图(video-edit) 建一颗 M6×30 螺栓导出 STEP(freecad) @geometry.png 做成中文讲题视频(edu-agent) --- 9. 和 Hermes 的关系(实用判断) 兼容路径 • 说明: Hermes 有 native MCP:可按 opencode 方式注册 uvx MCP + skill 价值最高 • 说明: core(读图/PDF/视频)+ video-memory + omni-av 注意 • 说明: 依赖 DashScope;部分工具要 ffmpeg;Blender/FreeCAD 要本机 App 重叠 • 说明: Hermes 已有 vision / youtube / video 等 skill;core 的 visualize/动态分辨率读文档更全 若要接到当前 Hermes,最小路径大概是: 1. 装 uv + ffmpeg 2. 配 DASHSCOPE_API_KEY 到 ~/.qwen-mm-plugins/config 3. 在 Hermes config.yaml 加 MCP server:qwen-mm-plugins-core 4. 把对应 skill/ 链到 Hermes skills 目录 --- 10. 学习路径建议 1. 先读 README + docs/en/installation.md + cookbooks/core/usage.md 2. 跑通 core:read_image → visualize(PDF) → ocr/grounding 3. 看 shared/: / / api_dashscope.py / mcp_framework.py 4. 再看一个生成/长视频能力:video-edit 或 video-memory 5. 扩展:照 capabilities/example + docs/how_to_add_new_capability.md 加自己的 cap --- 11. 一句话总结 Qwen-MM-Plugins = 跨 harness 的多模态“外挂操作系统”:用 Skill 教模型用工具,用 MCP+uvx 交付工具,用 DashScope/本地工具补齐读、搜、剪、建、讲全链路。 --- 要继续的话我可以直接帮你做其中一件: 1. 接到 Hermes(写 MCP + skill 配置) 2. 本机装 core 并 verify 3. 深挖某个能力源码(core / video-memory / edu-agent) (2/2)

Profilbild von Rishi
Rishivor 1 Monat

Multimodal agents become much more useful when perception tools are composable inside the harness, not bolted on as separate apps. The key test will be provenance: can each visual action show which frame, region, or document span drove it?

Ähnliche Videos

Explore state-of-the-art multimodal prompting in our new short course Large Multimodal Model Prompting with Gemini, taught by Erwin Huizenga in collaboration with Google Cloud. One interesting insight from this course: with multimodal models, prompt structure matters significantly. Placing text inputs, such as a patient's medical history, before image inputs, like an X-ray, can enhance the model's ability to contextualize and interpret visual data effectively. In other contexts, such as image captioning, you may get better results by putting the image first. Multimodal models behave differently than text-only LLMs, and effective prompting for models varies depending on the model you’re using. In this course you’ll learn how to effectively prompt Gemini models. Gemini's multimodal capabilities also enable new approaches in AI application development, for example: - The Gemini library handles various video formats (MP4, MOV, MPEG), streamlining applications using these formats. - Large context window (up to 1 million tokens) enables processing of extensive content, like analyzing multiple 50-minute videos simultaneously. - Function calling feature integrates real-time data (e.g., current exchange rates) into model responses. The course demonstrates building multimodal applications with real-world examples including document analyzers that reason across text and graphs simultaneously, video content extractors that find and timestamp specific information from multiple hours of footage, and automated expense report systems processing receipt images while cross-referencing company policies. Sign up here:

Andrew Ng

74,060 Aufrufe • vor 2 Jahren