Introducing MiniMax H3, our next-gen open-weight multimodal video model,... built for general intelligence beyond single-task generation. H3 understands text, images, video, and audio together, interpreting motion, sound, emotion, and cinematography as one unified creative language, delivering generation with fine-grained precision and cinematic quality. → Native Multimodal Understanding & Generation → Precision Editing & Control → Commercial-Grade Output: film, ads, MVs, UI, game CG → Cinematic Audio-Visual Quality: native stereo audio, up to 2K/24FPS Multi-asset reference for characters, motion, camera movement + voice cloning. Hailuo AI (MiniMax)'s H3 takes ideas from concept to production in one flow. Try it now: #MiniMaxH3show more

SARAH
97,618 views • 1 month ago
MiniMax H3 (Hailuo-03) is now available on fal. -... Open-weight multimodal video model with native stereo audio on every generation. - Combines up to 9 images, 3 video clips, and 3 audio clips as references - Holds subjects, motion, and audio consistent - Weights will be released soon.show more

fal
54,627 views • 1 month ago
MiniMax H3 is now 50% OFF on Magnific for... 2K video, only until September 1. 🔥 I’ve been trying MiniMax H3 on Magnific, and it feels like a big upgrade for AI video creation. It’s not just about turning text into videos. You can use text, images, videos, and audio together in one prompt, giving you more control over the final video. Here’s what makes it stand out: - Multimodal: Use text, images, video, and audio in one prompt. - Multiple references: Add up to 9 images, 3 videos, and 3 audio files. - 2K video: Create videos up to 15 seconds long. - Built-in sound: Generate voice, music, and sound effects with the video. - Easy editing: Remove objects or transfer motion easily. - More control: Control the camera, characters, and voice. You can use it to turn posters into videos, moodboards into short films, and product images into ads. It also helps bring your ideas to life with realistic movement, lighting, reflections, and sound. The workflow is simple: give it your references → generate → edit → refine. Try MiniMax H3 on Magnific:show more

Markandey Sharma
96,964 views • 18 days ago
MiniMax H3 is a serious upgrade for AI video... creation. Instead of relying on just text prompts, you can combine images, videos, and audio references to control the character, motion, camera, and sound. What stands out: → Up to 9 image + 3 video + 3 audio references → Native synced voice, music & sound effects → 2K video generation up to 15 seconds → Instruction-based video editing → Motion transfer → Better character, camera & voice control The interesting part is that you can build a full scene from references instead of endlessly regenerating until the result feels right. 🔗 Now is a great time to try MiniMax H3 on Magnific, especially with the current 2K offer.show more

Tanvir Anjum
32,271 views • 17 days ago
The next generation of AI video has arrived. Seedance... 2.5 combines long-form video generation, native audio, and precise scene control in a single model. — Text-to-video & image-to-video — Up to 30-second videos — Native synchronized audio — First & last frame guidance — Built for cinematic content Try now →show more

Wiro AI — Ship AI Faster
13,973 views • 26 days ago
MiniMax H3 is now on Magnific and honestly, there’s... a lot you can do with it. You can mix text, images, videos and audio in a single prompt up to 9 images, 3 videos and 3 audio references. Start creating now: It can generate up to 15s of 2K video with synced sound, including voice, music and effects. And you can go beyond generation too: edit clips, remove objects, transfer motion, and control the camera, character and voice. But the multi-reference workflow is probably my favorite. Give it your product, character, environment and motion references, and H3 pulls everything together. It feels like a much easier way to go from an idea to an actual finished video.show more

Kalsoom (ghotai )
47,888 views • 13 days ago
Seedance 2.0 is coming to Elser AI 🎬 Early... access → This isn’t just an update. It’s a studio-grade AI film engine upgrade. • Multi-shot storytelling • Long-form character consistency • Native audio-visual generation • Multimodal input + in-model editing • Advanced motion synthesis Built for real cinematic AI filmmaking. Big shift 👇 UNLIMITED 2-Year Creator Plan — 87% OFF One subscription. Seedance 2.0 + Kling 3 + Nano Banana Pro + Veo 3.1. If you’re serious about AI anime & cinematic storytelling — this changes the game. Early access →show more

Mujeeb Ahmed
17,094 views • 6 months ago
Seedance 2.0 is now live on invideo for Max,... Generative, and Team plan users. BytePlus's most advanced video model — and arguably the most controllable AI video model ever released. Multimodal input. Motion replication from reference videos. Native audio-video generation in one pass. Character consistency that actually holds. Director-level camera control. Real-world physics. This isn't prompt-and-pray. This is a production tool. Only available through business email verification for all regions except US and Japan.show more

Invideo
5,552,011 views • 5 months ago
Vidu Q3 is now available on Akool Experience next-level... AI video creation with Vidu Q3 — a powerful multi-modal model that seamlessly blends high-fidelity visuals with perfectly synchronized audio. Powered by Vidu AI, Q3 shines in: Deep narrative understanding Mastery of complex cinematic language Lifelike, human-feeling motion and emotion If storytelling and cinematic quality matter to you, Vidu Q3 is built for it. Try it now on Akool.show more

Akool Inc
1,528,592 views • 7 months ago
Dreamina Seedance 2.0 is Officially here! Dreamina Seedance 2.0,... ByteDance’s AI-powered creative platform, lets creators easily transform ideas into high-quality videos. You can now edit videos like images, using up to 4 reference modalities (video, image, audio, and text) with precise control over visual effects, camera movements, and more. Why is this revolutionary? - 🔥 One-Prompt Video Editing: Edit videos seamlessly as easily as editing images. - 🎬 Multimodal Creativity: Combine images, video, audio, and text—up to 12 files at once. - 🎥 Remake Viral Content: Create high quality, professional-level videos with intelligent AI. - 🔧 Full Creative Control: Keep consistency in shots, typography, camera flow, and more. This isn’t just another video editing tool, it’s a one-stop AI workspace for all your creative needs. Ready to level up your content? Explore Dreamina Seedance 2.0 and start creating today. 🔗 #dreamina #seedance2 #seedream5 #dreaminatutorial #ai #aitools #aidesign #ecommercedesign #digitalmarketing #startupbusinessshow more

GitHub Projects Community
18,440 views • 5 months ago
This AI just turned me into a film director…... No editing skills. No timeline headaches. Just one prompt. This is Seedance 2.0 🎬 You can literally combine: → Text → Images → Videos → Audio And it understands everything. Even crazier? You can control it like this: Image → character Video → camera movement audio1 → music/voice It doesn’t just generate clips… It builds full cinematic scenes with: → Consistent characters → Smooth transitions → Realistic motion → Built-in lip sync Basically… From a single prompt → you get a multi-shot story. Not AI video. AI filmmaking. Go try it before everyone catches on 👇show more

Kshitij Mishra | AI & Tech
60,393 views • 4 months ago
Introducing Muse Image and Muse Video, the first media... generation models developed by Meta Superintelligence Labs. Muse Image is our most advanced image generation model yet. It follows instructions faithfully, edits with precision, composes from multiple references, and draws on Instagram for social context. It also brings agentic tool use capabilities to image generation and integrates with Muse Spark. You can try Muse Image in the Meta AI app and web, as well as in Instagram Stories and WhatsApp – starting in limited countries with more locations on the way. Today we’re also previewing Muse Video, which is built upon the same pretraining base as Muse Image to deliver exceptional visual fidelity with native audio support. Learn more about both models:show more

AI at Meta
848,991 views • 1 month ago
currently experimenting with WAN 2.6 I2V on GMI Cloud... in this test, I’m comparing two audio workflows and honestly both perform really well. one scene uses audio generated directly from the prompt, while the other uses manually uploaded audio taken from the film 300. visually, both deliver strong motion and solid performance. however, the version with audio coming straight from the prompt feels slightly more refined, camera movement is smoother, transitions flow more naturally and the sync between voice, facial motion, and pacing feels more cohesive. lip sync, especially for Chinese dialogue also comes across a bit cleaner. you can choose single shot for a clean, focused moment or multi shot if you want more cinematic transitions, even when working from just one reference image. one important note: always turn on prompt extension. it makes a noticeable difference in how well the model understands motion, transitions and overall scene flow. both audio approaches are totally usable, but if you’re building dialogue driven or cinematic scenes, starting with audio from the prompt gives WAN 2.6 a bit more context to work with. I’ll be pushing this further with more dynamic camera movement and transitions next. more experiments coming soon ✨ Wanshow more

DStudioproject
97,086 views • 8 months ago
llama.cpp isn't just for text LLMs anymore. Pure C++... zero shot voice cloning just officially landed in mainline. Text generation was only step one. If you’re building autonomous local AI agents, real time voice assistants, or edge workflows, instant low latency audio is the missing piece. Thanks to PR #26254, Alibaba’s state of the art Qwen3 TTS model family is now natively supported directly inside the llama.cpp repository under the multimodal (mtmd) framework. No Python bloat. No massive PyTorch CUDA overhead. Just raw, hyper optimized C++ running GGUF voice weights. Here is why this native update is a massive deal for the open source local AI stack: # Multimodal Architecture (.gguf + mmproj) Qwen3-TTS splits the workload between the base language model backbone and a multimodal projection adapter. llama.cpp handles this using the llama-tts binary, mapping the text model alongside its --mmproj projector to process audio tokens seamlessly. # Zero Shot Voice Cloning in Seconds You don't need fine tuning or massive dataset training. Feed the C++ engine a single 5 to 10 second .wav audio sample using the --tts-speaker-file flag, and it accurately clones the exact timbre, tone, and accent on the fly. # Real World T4 GPU Benchmark & Resource FootprintRunning the 1.7B Base model in 8-bit quantization (Q8_0): - VRAM Footprint: ~7 GB peak VRAM during active zero-shot cloning. - Audio Quality: Studio grade, natural-sounding voice output in seconds. • - Execution: Direct execution via native compiled binaries or sub process calls. # Coming Next to llama-server (PR #26603) Beyond CLI execution, a native POST /tts HTTP endpoint is currently being added to llama-server, which will soon allow you to trigger voice generation directly via standard REST API requests! # quick note on Colab compilation: Because this code was merged into mainline very recently, pre-built third-party binaries haven't fully caught up yet. Compiling llama-tts directly from source on Google Colab's free CPU instance can take about 1 hour (or ~1-2 minutes if targeting single GPU arch like -DCMAKE_CUDA_ARCHITECTURES=75). Be patient during the build step, or compile it locally on your own rig for instant execution! To test this out yourself, I built a zero config Google Colab notebook that compiles llama.cpp, downloads the Q8_0 GGUF files from HuggingFace, and spins up an interactive Gradio Studio UI so you can record/upload 3 second clips and clone voices in real time. Stop sleeping on native C++ audio. The era of bulky Python audio pipelines is officially over. Links to the free Google Colab notebook and the official ggml org GGUF HuggingFace model repository are in the replies below! available in q4 and q8 both variants, 1 GB and 1.85 GBs respectively (requires additional ~500MB mmproj gguf) Are you building local voice agents yet? What does your current audio stack look like? Drop your setups below!show more

Alok
47,881 views • 28 days ago
Big moment for text-to-speech. Qwen just open-sourced a text-to-speech... model that lets you clone voices, design new ones, and control speech using natural language. Let me explain what I mean: You can literally tell it "speak in a cheerful tone with slight nervousness," and it actually does that. No complex audio engineering needed. What makes this special: - 3-second voice cloning - Covers 10 languages: English, German, French, and more - Latency as low as 97ms for real-time applications - Supports both streaming and non-streaming generation The model comes in two sizes (0.6B and 1.7B parameters), so you can pick based on your hardware and quality needs. Three modes to work with: 1. Custom Voice: Use pre-built premium voices with instruction-based style control 2. Voice Design: Describe the voice you want in plain English (or Chinese), and the model creates it 3. Voice Clone: Provide a 3-second reference audio and clone that voice The best part? It integrates with vLLM for production deployment and has a simple Python package you can pip install. I've shared a link to the GitHub repo in the next tweet.show more

Akshay 🚀
31,249 views • 7 months ago
Kling 3.0 for AI UGC videos is absolutely insane... 🤯 I spent 25,000+ credits in Kling perfecting the ultimate prompting framework to get the best AI UGC outputs possible. And the 3.0 update just made everything even better. Perfect for DTC brands and agencies who want high-quality AI UGC without paying $500/video for actual UGC. Here's what Kling 3.0 unlocks for AI UGCL → "AI Director" system that understands full scripts and auto-schedules camera angles (shot/reverse shot) in one generation → 3 to 15 second cinematic clips with full temporal coherence → Improved character and element locking so your subject stays consistent across shots and angles → Native 4K output for both video and stills — actually usable for professional ad creative And best of all: perfect character consistency ACROSS different shots. What this means for AI UGC: - Multi-shot storytelling in a single generation cycle - Longer clips that actually hold together - Consistent characters across your entire ad - Output quality that's ready for paid media This is the closest AI video has gotten to replacing a real shoot for performance creative. I recorded a full breakdown of the prompting framework I built after burning through 25,000 credits. Want access to all the prompts I use to create AI UGC with Kling? > Like this post > Comment "KLING" And I'll send it over (must be following so I can DM)show more

Mike Futia
25,567 views • 6 months ago
We are in an insane run of open-weight drops.... Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: 🧠 LLMs & Reasoning → DeepSeek-V4-Flash-0731 (my king 👑): 304B MoE refresh, Terminal-Bench 2.1 jumps 61.8→82.7 over the preview, DeepSWE 7.3→54.4. Closes in on Opus-4.8 on Agents' Last Exam (25.2 vs 25.7). MIT. → Muse-Glimmer-30B, from Meta (they are back!!): their first open agentic model. ~29.6B dense + perception encoder, 131k+ context, built to run fully local, no cloud. Apache 2.0. → Liquid AI LFM2.5-2.6B: 2.69B params, 131k context, 220 tok/s on an M5 Max in under 2.5GB RAM. Competitive with models 4x larger on agentic tasks. → inclusionAI Ling-3.0-flash: 124B total, only 5.1B active, ~12% the size of their old 1T flagship Ring-2.6, matches it on key benchmarks. MIT. → inclusionAI Ling-3.0-tiny: 7.9B total, 1.3B active, 86-90 tok/s on an M4 Pro MacBook at ~8GB peak memory. MIT. → NVIDIA Nemotron-3.5-Lightning-30B-A3B: hybrid Mamba-2+MoE+Attention, up to 1M context, runs on a single H100 or DGX Spark, SWE-bench Verified 52.8. → deepgrove maple-preview: 20B-A1B ternary-weight reasoner, 218 tok/s on a Mac mini M4, 5.3GB checkpoint. MIT. → BigBang-v1 (endless-frontier): fine-tuned from Qwen3.6-35B-A3B via a self-evolving generator/critic synthetic-data loop. Lands aggregate performance between DeepSeek V4 Flash (284B) and V4 Pro (1.6T), at 35B. Apache 2.0. 🎬 Video → MiniMax-H3: 33B dense omni model, native stereo audio, up to 2K/15s. 3.6k+ likes already. → Minimax-H3-Turbo (lightx2v): Apache-2.0 turbo distillation of H3 for fast inference. → Lightricks LTX-2.5: image-to-video update, custom Gemma-4-12B text encoder, a markedly stronger distilled model. 🔊 Voice → NVIDIA NemotronLabs VoiceChat-11B: full-duplex speech-to-speech, ~450ms turn-taking, #2 on open VoiceBench, and the first open full-duplex model with live tool-calling mid-conversation. 🛡️ Safety → Mistral Shieldstral-1.0-3B: 3B multimodal guardrail that takes your safety policy as plain text instead of fixed categories. Beats LlamaGuard-4-12B and ShieldGemma-9B on HarmBench (99.4) and ToxicChat (84.1) at a fraction of the size. Apache 2.0.show more

Victor M
54,264 views • 21 days ago
I’ve used all the recent GenAI video models extensively... & here’s my 2¢: 🎬 Runway Gen3 Alpha - best image quality & motion for text-to-video & embedded words. Great at prompt travel changes over the course of 10 sec. And I’m super bullish on how gen3 will evolve, hopefully adopting the features listed below. Kling - best quality for image-to-video with prompt control, like eating food. Great clip extension that accounts for character (ie walking stride) & camera movement (speed & angle), rather than just using final frame. But it’s limited availability & Chinese native language is limiting. Used for Spider-Man video below (via Midjourney). LumaLabs - best for keyframe start & end control (it can not be overstated how important this is. other services should add it ASAP!) and their high dynamic action movements are really fun. Luma was used in my viral Multiverse of Memes video. PikaLabs - they haven’t gotten as much attention as others lately. But they did update their video model a few weeks ago and it looks great. Also, they are notable for their unique & AWESOME features, like video in-painting & out-painting. My perfect AI video platform would have the following features: 1) Gen3’s quality, prompt control & text embedding. 2) KLing’s image-to-video quality, prompt control & clip extension quality. 3) Luma’s multi-keyframe control & dynamic movement ability. 4) Pika’s inpainting & outpainting ability. And a video-to-video (aka next-gen Runway gen1) could be a game changer, too. It’s an exciting time to be alive 🫶 Who will get there first? 🔉🔉show more

Blaine Brown
26,535 views • 2 years ago
You don't need a GPU for fast studio grade... voice cloning anymore. Qwen3 TTS (1.7B Q4_K_M) + mainline llama.cpp is officially the fastest way to generate zero shot voice clones using 100% pure CPU execution. Following up on my last post where we ran the Q8 model on a GPU, we just took local C++ voice synthesis a massive step further. The open source community quantized Alibaba's SOTA Qwen3 TTS model down to Q4_K_M GGUF, completely freeing local audio pipelines from dedicated graphics hardware. Here is the real world benchmark and hardware breakdown of running SOTA voice cloning on CPU: # Architecture & Model Setup Using Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf paired with the 8 bit multimodal projector (mmproj-Q8_0.gguf), llama.cpp executes the entire pipeline in pure C++. No PyTorch, no CUDA dependencies, and no VRAM bottlenecks. # Real-World Memory Footprint - Baseline RAM: 1.6 GB system idle. - Peak Generation RAM: 8 GB RAM during active voice synthesis. - Requirement: Any basic machine with at least 8 GB of system RAM can run this easily. # Real World CPU Benchmarks - Google Colab Free Tier (Throttled 2 Core CPU): Synthesizes a 5 sec studio quality audio clip (~8 words) in 45 seconds. - Modern Consumer CPU (Intel i5/i7 13th/14th Gen or AMD Ryzen 7000/9000): generation should drop to 5 to 20 seconds (nearly 1:1 real-time generation speed!). # Zero Shot Voice Cloning Quality Pass any 5 to 20 second .wav audio sample to the C++ engine using the --tts-speaker-file flag. It yields clean, natural sounding cloned speech with virtually zero quality loss compared to unquantized FP16 weights. To make testing seamless, I built an updated zero config Google Colab notebook. It pulls the official pre built llama.cpp CPU binaries (zero compilation time!) launches a live Gradio web app right in your browser. Record a 5 second clip from your mic (or drop a .mp3, .wav file), type text, and generate cloned audio on CPU. Native C++ audio models are making edge based, offline AI voice agents a reality. Links to the free Q4 CPU Colab notebook and the Q4_K_M GGUF HuggingFace repository are in the replies below! Which models have you been running on your CPUs? What CPU hardware are you using for local inference?show more

Alok
60,514 views • 25 days ago
Google dropped a new AI paper called LUMIERE. It's... remarkably flexible, supporting video inpainting, image-to-video, AND stylized video generation tasks. Say hello to “space-time diffusion” for video generation! Now what the heck does that mean exactly?! 🌐⏳ → TL;DR it utilizes a “Space-Time UNet” architecture that generates the full duration of the video in one pass, rather than generating distant keyframes and interpolating between them like prior works. Because the computation is done in this “compressed space-time representation” to generate the full clip at once, it's far more temporally consistent. → Another benefit of generating the full video at once is that you can “direct” the video generation, making it easier to hand off to other models/tasks without having to stitch together partial solutions. You can condition generations on additional inputs, meaning you get the full stack of AI video capabilities – from video inpainting to image-to-video and beyond. → New SOTA for AI video generation? User study results in the paper suggest human evaluators preferred Lumiere over Runway Gen-2, Pika Labs, and Stable Video Diffusion in terms of quality, text alignment AND motion. But as always, we need to get hands-on with this tech when Google *actually* decides to ship it. → Could this end up inside YouTube? Y’all know i’m obsessed with blending reality and imagination – so it’s the video inpainting tech I'm most excited about. I really hope this model finds its way into YouTube's Generative AI efforts, and based on their prior announcements and the list of acknowledgments in the paper I think it might! 🤞🏽 Links: 🔗Paper: 🔗Project:show more

Bilawal Sidhu
44,822 views • 2 years ago