Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

👀 Seeing is just the beginning. With Qwen-MM-Plugins, turn your favorite agent harness multimodal-native — read images, videos & documents, edit videos, work with 3D/CAD, and more. From multimodal models → multimodal agents. 🚀 Watch it in action:

796,350 görüntüleme • 1 ay önce •via X (Twitter)

32 Yorum

RomboDawg profil fotoğrafı
RomboDawg1 ay önce

The entire open source AI community waiting for Qwen-3.8-27b announcement

Guillermo profil fotoğrafı
Guillermo1 ay önce

Hi @Alibaba_Qwen, You may find this useful: It's a proof-of-concept to bring coding agents into Office via ACP. Here is a demo presentation created by Qwen 3.8 Max using OpenCode:

Mert Kaya profil fotoğrafı
Mert Kaya1 ay önce

Sadly I do think it works only through api, am I missing something here? Supporter of Qwen and their approach on AI but this? Also Claude Code as a contributor is funny.

Cyril Dieumegard 🇨🇭 profil fotoğrafı
Cyril Dieumegard 🇨🇭1 ay önce

The 3D/CAD part is what catches my eye. That’s a much more interesting bridge than just reading another PDF.

Hex Nova profil fotoğrafı
Hex Nova1 ay önce

This is actually huge 🔥 Finally agents that can properly handle videos, PDFs and even 3D without all the usual frame-by-frame or page-by-page headaches. Qwen-MM-Plugins looking really promising. Gonna try this out today 🚀

altermouv profil fotoğrafı
altermouv1 ay önce

i can't handle all these new releases.. we have a multimodal plugins now! so there will be no Qwen 3.8-omni version?

liemfukliang profil fotoğrafı
liemfukliang1 ay önce

I hope the QwenLM can be used with Qwen 3.8 27B / 35B and can make ai video compilation. I just add all my video from the camera, the ai analyse and understand all the video. Then the ai create a video compilation based on the prompt.

VenkyBeast profil fotoğrafı
VenkyBeast1 ay önce

Qwen 3.8 27B when!? We are still waiting 😭

Mr Alexander Official profil fotoğrafı
Mr Alexander Official1 ay önce

This is where multimodal agents start feeling genuinely useful — not just seeing content, but actually acting on it. Reading videos and documents, editing them, and even working with 3D/CAD opens up a much bigger agent workflow. 🚀

TinyHumans AI profil fotoğrafı
TinyHumans AI1 ay önce

excited for this! great work

Maurice | KI-Vater profil fotoğrafı
Maurice | KI-Vater1 ay önce

This is the exact moment open-source agents stopped playing catch-up. Multimodal-native out of the box. Video edit. 3D/CAD. Long-video memory. Qwen cooked again. 🚀

Amiee profil fotoğrafı
Amiee1 ay önce

@Alibaba_Qwen Important correction: my 10.69M-token measurement was during the 50% off-peak window. Normal rates: $6 ≈22.91M/30d $18 ≈91.66M/30d 45.83M / 183.32M are best-case off-peak numbers. Please respond to this.

Jayadeep Reddy profil fotoğrafı
Jayadeep Reddy1 ay önce

The 3D/CAD and video-editing plugin hooks are the real signal here. Most enterprises do not need a better chatbot. They need an agent that can manipulate their existing design and media pipelines without rewriting them.

Rey Neill profil fotoğrafı
Rey Neill1 ay önce

I’m surprised this idea has been untouched for so long

Linen profil fotoğrafı
Linen1 ay önce

I think people still underestimate how important harnesses are becoming. At some point, the question won’t just be “which model is best?” It’ll be “which model + tools + workflow gives the best result?” Those plugins seems a good step in that direction. What’s the first plugin/tool you’d want in your agent setup?

Sarcastic Badger profil fotoğrafı
Sarcastic Badger1 ay önce

I like that it’s s giving agents native access to the same visual context humans use, rather than treating images and documents as separate tools.

AI Mastery Guide profil fotoğrafı
AI Mastery Guide1 ay önce

Multimodal native agents that handle 3D and CAD too, that's a big jump.

fafa.👩🏻‍💻 profil fotoğrafı
fafa.👩🏻‍💻1 ay önce

love it!

Aivora profil fotoğrafı
Aivora1 ay önce

Honestly impressive for a free generation, the motion and lighting nailed it. The one thing giving it away is the resolution, free tier just doesn't have the crispness the paid version probably delivers.🤑

elliott profil fotoğrafı
elliott1 ay önce

@grok what subs or API service keys do I need to utilize something like this? can this be used to add multimodal capabilities to say - the Reasonix CLI that runs DeepSeek? Or is this focused on QOder type harnesses?

OwlcoreAI profil fotoğrafı
OwlcoreAI1 ay önce

3D/CAD is getting pretty popular recently. Interesting to see what the cost will be for this.

泰多宝 profil fotoğrafı
泰多宝1 ay önce

the plugin route is the pragmatic call. most agent harnesses are still text-native with multimodal bolted on through file paths, so making images and video first-class citizens of the loop matters more than another point release of the base model

ShadowAguy profil fotoğrafı
ShadowAguy1 ay önce

seeing is just the beginning, but reading the docs is the real adventure

WuBu ⪋ WaefreBeorn 🇺🇸 👑 profil fotoğrafı
WuBu ⪋ WaefreBeorn 🇺🇸 👑1 ay önce

very nicr

Elmer den Braven profil fotoğrafı
Elmer den Braven1 ay önce

This is a big jump in capability!

Ahmed 🇵🇸🍉 🔻 profil fotoğrafı
Ahmed 🇵🇸🍉 🔻1 ay önce

hey @FreeCADNews you had an appearance here ❤️

PEF profil fotoğrafı
PEF1 ay önce

Je viens de les installer pour Claude Code. C'est vraiment génial @Alibaba_Qwen.

John profil fotoğrafı
John1 ay önce

Open source it

AgentX profil fotoğrafı
AgentX1 ay önce

👀📝

Ray profil fotoğrafı
Ray1 ay önce

@grok can I do this on my Msi 1050 8gen ?

老林学KOPI AI AGENT profil fotoğrafı
老林学KOPI AI AGENT1 ay önce

-4k.png 读出每一个数字 @report.pdf 总结第 3 页 @receipt.jpg OCR 并加总金额 @street.jpg 把每辆车框出来 @lecture-2h.mp4 主要观点 + 时间戳(video-memory) @meeting.mp4 带说话人时间戳转写(omni-av) 生成一张小熊猫深夜敲代码的图(video-edit) 建一颗 M6×30 螺栓导出 STEP(freecad) @geometry.png 做成中文讲题视频(edu-agent) --- 9. 和 Hermes 的关系(实用判断) 兼容路径 • 说明: Hermes 有 native MCP:可按 opencode 方式注册 uvx MCP + skill 价值最高 • 说明: core(读图/PDF/视频)+ video-memory + omni-av 注意 • 说明: 依赖 DashScope;部分工具要 ffmpeg;Blender/FreeCAD 要本机 App 重叠 • 说明: Hermes 已有 vision / youtube / video 等 skill;core 的 visualize/动态分辨率读文档更全 若要接到当前 Hermes,最小路径大概是: 1. 装 uv + ffmpeg 2. 配 DASHSCOPE_API_KEY 到 ~/.qwen-mm-plugins/config 3. 在 Hermes config.yaml 加 MCP server:qwen-mm-plugins-core 4. 把对应 skill/ 链到 Hermes skills 目录 --- 10. 学习路径建议 1. 先读 README + docs/en/installation.md + cookbooks/core/usage.md 2. 跑通 core:read_image → visualize(PDF) → ocr/grounding 3. 看 shared/: / / api_dashscope.py / mcp_framework.py 4. 再看一个生成/长视频能力:video-edit 或 video-memory 5. 扩展:照 capabilities/example + docs/how_to_add_new_capability.md 加自己的 cap --- 11. 一句话总结 Qwen-MM-Plugins = 跨 harness 的多模态“外挂操作系统”:用 Skill 教模型用工具,用 MCP+uvx 交付工具,用 DashScope/本地工具补齐读、搜、剪、建、讲全链路。 --- 要继续的话我可以直接帮你做其中一件: 1. 接到 Hermes(写 MCP + skill 配置) 2. 本机装 core 并 verify 3. 深挖某个能力源码(core / video-memory / edu-agent) (2/2)

Rishi profil fotoğrafı
Rishi1 ay önce

Multimodal agents become much more useful when perception tools are composable inside the harness, not bolted on as separate apps. The key test will be provenance: can each visual action show which frame, region, or document span drove it?

Benzer Videolar

Explore state-of-the-art multimodal prompting in our new short course Large Multimodal Model Prompting with Gemini, taught by Erwin Huizenga in collaboration with Google Cloud. One interesting insight from this course: with multimodal models, prompt structure matters significantly. Placing text inputs, such as a patient's medical history, before image inputs, like an X-ray, can enhance the model's ability to contextualize and interpret visual data effectively. In other contexts, such as image captioning, you may get better results by putting the image first. Multimodal models behave differently than text-only LLMs, and effective prompting for models varies depending on the model you’re using. In this course you’ll learn how to effectively prompt Gemini models. Gemini's multimodal capabilities also enable new approaches in AI application development, for example: - The Gemini library handles various video formats (MP4, MOV, MPEG), streamlining applications using these formats. - Large context window (up to 1 million tokens) enables processing of extensive content, like analyzing multiple 50-minute videos simultaneously. - Function calling feature integrates real-time data (e.g., current exchange rates) into model responses. The course demonstrates building multimodal applications with real-world examples including document analyzers that reason across text and graphs simultaneously, video content extractors that find and timestamp specific information from multiple hours of footage, and automated expense report systems processing receipt images while cross-referencing company policies. Sign up here:

Andrew Ng

74,060 görüntüleme • 2 yıl önce