🚀 LongCat-Video Now Open-Source: Text/Image-to-Video + Video Continuation in... One Model 🏆 Text/Image-to-Video Performance Hits Open-Source SOTA 🎬 Minutes-Long High-Quality Videos: No Color Drift/Quality Loss (Industry-Standout) ⚙ 13.6B Params | Strong Open-Source DiT-Based Unified Multitask Video Base Model ⚡ C2F Pipeline + Block Sparse Attention: 720p/30fps Video in Minutes 🤗 Open-Source Links: GitHub: Hugging Face: Project Page:show more

Meituan LongCat
43,802 просмотров • 9 месяцев назад
We are building “Open Source Nano Banana for Video”... - here is open source demo v0.1 We are open sourcing Lucy Edit, the first foundation model for text-guided video editing! Get the model on Hugging Face 🤗, API on @FAL, and nodes on ComfyUI 🧵show more

Decart
414,301 просмотров • 10 месяцев назад
LTX-2 is THE open source video AI moment I've... been waiting for--the first to generates high quality video AND is FAST!! - 4090: Generate 20s at 720p in 2 minutes - A4500 (3070): Generate 10s at 480p in 3 minutes 1-click install on Pinokio, do it now!show more

cocktail peanut
116,072 просмотров • 6 месяцев назад
Stanford dropped FramePack This AI can run on 6... GB laptop GPU to generate minute long 30fps video from single image No distillation, open source. 10 wild examples & how to try it: 👇show more

Min Choi
633,607 просмотров • 1 год назад
🚨BREAKING: The best free AI video tool just dropped.... Now you can generate high-quality 4K videos just by using a text description. SkyReels V2 is open source, unlimited, and completely free. Here’s how it works (with real examples):👇show more

Hasan Toor
254,320 просмотров • 1 год назад
Woke up today to see spacepxl trained a model... on top of Wan to achieve SOTA performance in deblurring w/ an approach that will generalise to all kinds of video control tasks - upscaling, canny, video inpainting, etc. I love the open source AI art/banodoco community so muchshow more

POM
45,609 просмотров • 1 год назад
🚨 SkyReels just launched! The world’s first open-source video... generation platform supporting unlimited duration 🔥 All-in-one creator toolkit: - Consistent high-quality video (LoRA ready) - Fast gen, amazing output. - Amazing facial expressions . Plus text-to-film agent handles everything: script, character, storyboards, full AV gen, auto-edit . it's Wild! Step by step tutorial 👇show more

AshutoshShrivastava
81,012 просмотров • 1 год назад
🚨 BREAKING: Video generation just got its biggest upgrade.... WAN 2.6 generates complete audiovisual experiences in one pass—no stitching, no external tools, no manual sync. First open-source model to do this. Here are 5 examples:show more

Future Stacked
614,676 просмотров • 7 месяцев назад
🚨 BREAKING: Hollywood should be worried. AI just wiped... out the most expensive part of video production. No crew. No sound designer. No post-production pipeline. Alibaba’s open-source model WAN 2.6 generates complete audiovisual scenes in one shot. 7 insane examples 👀:show more

Jaynit Makwana
67,726 просмотров • 7 месяцев назад
LTX-2.3 is now live on OpenArt. 🎬 The most... capable open video model just got a major upgrade and you can use it right now. What's new in 2.3: → Sharper fine detail: Hair, textures, text, edges. All of it. → Tighter prompt adherence: Complex multi-subject prompts? Handle it. → Stronger image-to-video: less freezing, less Ken Burns drift, more actual motion. → Cleaner audio: fewer artifacts, tighter sync across text-to-video and audio workflows. → Native portrait: up to 1080×1920, trained on vertical data.show more

OpenArt
2,150,785 просмотров • 3 месяцев назад
Meet LongCat-Video-Avatar 1.5🐱—our upgraded, open-source digital human framework. Built... for real production, not just short demos. What's New: 🔹 Upgraded Audio Encoder: Replaces Wav2Vec2 with Whisper-Large, yielding significantly smoother and more natural lip dynamics. 🔹 Production-Ready Stability: Achieves accurate lip-synchronization, full-body temporal stability, and robust long-video generation with strict identity consistency. 🔹 Stylized Domain Generalization: Robustly generalizes to anime, animals, and complex real-world conditions such as multi-person interactions and object handling. 🔹 Efficient 8-Step Inference: Advanced step distillation accelerates inference to 8 NFE, balancing cost-effective serving with exceptional visual fidelity. 📊 LongCat-Video-Avatar 1.5 performs strongly in realism, naturalness, and stability, outperforming leading open-source models and closed systems. 🐱 Avatar 1.5 framework is now open source: 🔗 Weights & Code: 🔗 HuggingFace: 🔗 Tech Report: 🔗 Project Page:show more

Meituan LongCat
31,077 просмотров • 2 месяцев назад
Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation paper page:... Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring temporal consistency across video frames remains a formidable challenge. This paper proposes a novel zero-shot text-guided video-to-video translation framework to adapt image models to videos. The framework includes two parts: key frame translation and full video translation. The first part uses an adapted diffusion model to generate key frames, with hierarchical cross-frame constraints applied to enforce coherence in shapes, textures and colors. The second part propagates the key frames to other frames with temporal-aware patch matching and frame blending. Our framework achieves global style and local texture temporal consistency at a low cost (without re-training or optimization). The adaptation is compatible with existing image diffusion techniques, allowing our framework to take advantage of them, such as customizing a specific subject with LoRA, and introducing extra spatial guidance with ControlNet. Extensive experimental results demonstrate the effectiveness of our proposed framework over existing methods in rendering high-quality and temporally-coherent videos.show more

AK
375,123 просмотров • 3 лет назад
Before the week ends, let's acknowledge one of the... most INSANE week ever for open AI, with 25+ notable open-weight drops across every modality: 🧠 LLMs → NVIDIA Nemotron 3 Ultra: 550B hybrid Mamba-MoE, only 55B active, 1M context, MMLU 89.1. NVFP4 variant claims ~5x throughput on Blackwell. First openly-weighted 550B hybrid Mamba-Transformer, closing the gap with frontier closed models. → Google Gemma 4 12B: fully open dense any-to-any (text/image/audio/video), 256k context, encoder-free, 140+ languages, AIME 2026 at 77.5. Shipped with a 23-checkpoint QAT wave (mobile ONNX + MLX). Most deployable model of the week. → StepFun Step-3.7-Flash: 198B sparse MoE VLM, ~11B active, SWE-Bench PRO 56.3. Apache 2.0. → Liquid AI LFM2.5-8B-A1B: edge MoE, just 1.5B active, 128k ctx, MATH500 88.8, MLX-ready. Best on-device option this week. → JetBrains Mellum2-12B-A2.5B-Thinking: their first open MoE, near-Qwen3-14B coding at 2.5B active. Apache 2.0. 🎨 Image gen (the surprise of the week) → Ideogram 4: their FIRST-EVER open weights. 9.3B flow-matching DiT trained from scratch. #2 overall behind GPT Image 2, top open-weight model on Design Arena + LMArena. Strongest open checkpoint for text-rich images, full stop. It has taste. Still can't believe this is open weights. 🔊 Audio & Speech (a breakout week for open TTS, 4 labs shipped) → Boson Higgs Audio v3 4B: 102 languages, 21 emotions, singing/whispering/shouting, sub-second TTFA. → RedNote dots.tts: the only fully continuous (no codec) open TTS pipeline, Apache 2.0. → Google Magenta RealTime 2: real-time music gen, <200ms latency, text+audio+MIDI. multimodalart ported it to PyTorch within hours with live ZeroGPU demos. → NVIDIA Nemotron-3.5 ASR: 600M streaming, 17x more concurrent streams vs Parakeet RNNT 1.1B. 👁️ Vision & VLMs → PaddleOCR-VL-1.6: SOTA document parsing at 1B params, Apache 2.0. → Baidu NAVA: 6.3B joint audio-video gen, best-in-class A/V sync, Apache 2.0. 🎬 Video, 3D & World Models → NVIDIA Cosmos3-Super: 64B omnimodal world model coupling action trajectories with video+audio gen, for Physical AI. → JD JoyAI-Echo: up to 5-min multi-shot text-to-video on LTX-2.3. → ByteDance Bernini-R + VAST TripoSplat (single-image-to-3D Gaussian splats, MIT).show more

Victor M
539,883 просмотров • 1 месяц назад
(1/n) 🚀 With FastVideo, you can now generate a... 5-second video in 5 seconds on a single H200 GPU! Introducing FastWan series, a family of fast video generation models trained via a new recipe we term as “sparse distillation”, to speed up video denoising time by 70X! 🖥️ Live demo: (Thanks to @gmicloud for the support!) 🔗 Blog: 🔓 We fully open-source our models, code, and data with Apache-2.0 licensesshow more

Hao AI Lab
78,660 просмотров • 11 месяцев назад
30 minutes of video. Robot learns the task. Open-source,... end-to-end. An open-source framework for training robot policies from only 30 minutes of human egocentric videos captured via Meta Aria glasses: Achieving zero-shot transfer to robots without any robot data collection. The method relies on Interaction-Centric Tokens that encode hand-object spatial relationships invariant to embodiment and viewpoint, supplemented by auxiliary objectives like object motion prediction and latent consistency to extract richer supervision signals from the same data. HumanEgo demonstrates strong cross-embodiment, cross-environment performance on bimanual tasks, outperforming baselines like ACT and teleop data while being trainable on a single RTX 4090 GPU. Thanks for sharing, Zhi (Leo) Wang. 📌 Website: Paper: Code: Video: ——- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
17,077 просмотров • 2 месяцев назад
xAI isn't playing around. They just released the Grok... Imagine API, a unified video + image generation toolkit, and it's already sitting at #1 on the Artificial Analysis Video Arena for both Text-to-Video AND Image-to-Video. It's beating: ● Google's Veo 3.1 & Veo 3 ● OpenAI's Sora 2 ● Runway Gen-4.5 ● Kling 2.5 Turbo The Numbers Don't Lie: ● 64.1% win rate against Runway Aleph in blind human evaluations ● 57% win rate against Kling o1 ● Best-in-class latency. Sub-20 second generation for 720p, 8-second videos. (up to 15-second video) ● Native audio generation baked right into video output (dialogue, music, sound effects, all synced) What Makes It Different It's built for real creative workflows: ✅ Text-to-video AND image-to-video in one API ✅ Video editing with prompt-based controls (add/remove objects, restyle scenes) ✅ Camera controls: zoom, pan, timelapse, pull-back ✅ Style transfers: cyberpunk, watercolor, anime, you name it ✅ Performance animation: map your movements onto characters ✅ Native audio-video sync (no post-production needed) Why the focus on speed and cost? The partner feedback that shaped this: "Quality alone isn't enough if latency and cost make iteration painful." So xAI optimized for all three. Speed. Cost. Quality. Already Integrated With: ● fal. ai ● ComfyUI ● InVideo ● Flora ● HeyGen xAI went from underdog to chart-topper. The Grok Imagine API is fast, affordable, and genuinely production-ready. If you're building anything with AI video, this just became the one to beat.show more

tetsuo
18,325 просмотров • 5 месяцев назад
I’ve used all the recent GenAI video models extensively... & here’s my 2¢: 🎬 Runway Gen3 Alpha - best image quality & motion for text-to-video & embedded words. Great at prompt travel changes over the course of 10 sec. And I’m super bullish on how gen3 will evolve, hopefully adopting the features listed below. Kling - best quality for image-to-video with prompt control, like eating food. Great clip extension that accounts for character (ie walking stride) & camera movement (speed & angle), rather than just using final frame. But it’s limited availability & Chinese native language is limiting. Used for Spider-Man video below (via Midjourney). LumaLabs - best for keyframe start & end control (it can not be overstated how important this is. other services should add it ASAP!) and their high dynamic action movements are really fun. Luma was used in my viral Multiverse of Memes video. PikaLabs - they haven’t gotten as much attention as others lately. But they did update their video model a few weeks ago and it looks great. Also, they are notable for their unique & AWESOME features, like video in-painting & out-painting. My perfect AI video platform would have the following features: 1) Gen3’s quality, prompt control & text embedding. 2) KLing’s image-to-video quality, prompt control & clip extension quality. 3) Luma’s multi-keyframe control & dynamic movement ability. 4) Pika’s inpainting & outpainting ability. And a video-to-video (aka next-gen Runway gen1) could be a game changer, too. It’s an exciting time to be alive 🫶 Who will get there first? 🔉🔉show more

Blaine Brown
26,535 просмотров • 2 лет назад
SOMEONE GOT TIRED OF PAYING HIGGSFIELD AI'S SUBSCRIPTION SO... HE REBUILT THE WHOLE THING AND OPEN-SOURCED IT 200+ models. text-to-image, image-to-image, text-to-video, image-to-video all in one interface you configure a virtual camera in the Cinema Studio. pick the body, the lens, the focal length, the aperture and it writes the optimized cinematic prompt for you. completely in the background you never touch the camera keywords. you just set up the shot like a real cinematographer would Kling v3, Sora 2, Veo 3, Flux Dev, Midjourney v7, GPT-4o, Seedream 5.0, Runway Gen-3 all in there self-hosted. MIT licensed. runs on your machine. your data stays local the only thing you pay for is the model API calls themselves someone built this so you never have to pay Higgsfield AI againshow more

Rimsha Bhardwaj
101,058 просмотров • 3 месяцев назад
Dreamina Seedance 2.0 is Officially here! Dreamina Seedance 2.0,... ByteDance’s AI-powered creative platform, lets creators easily transform ideas into high-quality videos. You can now edit videos like images, using up to 4 reference modalities (video, image, audio, and text) with precise control over visual effects, camera movements, and more. Why is this revolutionary? - 🔥 One-Prompt Video Editing: Edit videos seamlessly as easily as editing images. - 🎬 Multimodal Creativity: Combine images, video, audio, and text—up to 12 files at once. - 🎥 Remake Viral Content: Create high quality, professional-level videos with intelligent AI. - 🔧 Full Creative Control: Keep consistency in shots, typography, camera flow, and more. This isn’t just another video editing tool, it’s a one-stop AI workspace for all your creative needs. Ready to level up your content? Explore Dreamina Seedance 2.0 and start creating today. 🔗 #dreamina #seedance2 #seedream5 #dreaminatutorial #ai #aitools #aidesign #ecommercedesign #digitalmarketing #startupbusinessshow more

GitHub Projects Community
18,440 просмотров • 4 месяцев назад
World Model is trending— let's revisit our HunyuanWorld journey.... We’ve been pioneering open-source 3D world generation in the past two months, and this ride’s only getting started. 🌍 📅 July: HunyuanWorld 1.0 📌 First open-source 3D world model compatible with CG pipelines (Unity/Unreal/Blender) 📌 Hit 2K+ GitHub stars in just two months ⭐—thank you for the love! 📅 August: 1.0-Lite 📌Same top-tier quality, running on consumer GPUs! 📅 September: 1.0-Voyager 📌 Direct 3D output + world memory—taking exploration further! Seamlessly integrated into CG pipelines with layered 3D modeling (assets, terrain, skybox) and fully open-sourced.. we’re fully committed to building open-source spatial intelligence for all! 🚀 💡 Why it matters? ✅ Seamless CG Pipeline Integration: Export generated 3D scenes as standard mesh formats, effortlessly integrating into industry-standard tools like Blender, Unity, and Unreal Engine for direct editing, animation, and physical simulation. ✅ Hierarchical Scene Editing: Deconstruct scenes into semantic layers (sky, background, foreground objects) via instance recognition and layer decomposition, allowing for atomic-level control—independently modify, relocate, or replace objects without rebuilding the entire world. Project page: Github: Amazing creations by Stijn Spanhove camenduru GENEL | AIを用いた動画制作 apolinario 🌐 とりにく Directive Creator 🪥 👇 #AI #3DGeneration #OpenSource #WorldModels #Hunyuan3D #HunyuanWorldshow more

Tencent HY
20,178 просмотров • 10 месяцев назад