Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

XAI EXTENDS GROK IMAGINE VIDEO GENERATION TO 10 SECONDS WITH QUALITY ENHANCEMENTS xAI has updated its Grok Imagine tool to produce videos lasting 10 seconds, doubling the prior limit. This change, along with refinements in visual and audio elements, expands the tool's utility for short-form content creation. xAI released...

28,881 görüntüleme • 7 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,265,716 görüntüleme • 1 yıl önce

Google presents Still-Moving Customized Video Generation without Customized Video Data Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation is still in its infancy, primarily due to the lack of customized video data. In this work, we introduce Still-Moving, a novel generic framework for customizing a text-to-video (T2V) model, without requiring any customized video data. The framework applies to the prominent T2V design where the video model is built over a text-to-image (T2I) model (e.g., via inflation). We assume access to a customized version of the T2I model, trained only on still image data (e.g., using DreamBooth or StyleDrop). Naively plugging in the weights of the customized T2I model into the T2V model often leads to significant artifacts or insufficient adherence to the customization data. To overcome this issue, we train lightweight Spatial Adapters that adjust the features produced by the injected T2I layers. Importantly, our adapters are trained on "frozen videos" (i.e., repeated images), constructed from image samples generated by the customized T2I model. This training is facilitated by a novel Motion Adapter module, which allows us to train on such static videos while preserving the motion prior of the video model. At test time, we remove the Motion Adapter modules and leave in only the trained Spatial Adapters. This restores the motion prior of the T2V model while adhering to the spatial prior of the customized T2I model. We demonstrate the effectiveness of our approach on diverse tasks including personalized, stylized, and conditional generation. In all evaluated scenarios, our method seamlessly integrates the spatial prior of the customized T2I model with a motion prior supplied by the T2V model.

AK

40,485 görüntüleme • 2 yıl önce

This is probably the most complex workflow I’ve ever built, only with open-source tools. It took my 4 days. It takes four inputs: author, title, and style; and generates a full visual animated story in one click in ComfyUI . I worked on it for four days. There are still some bugs, but here’s the first preview. Here’s a quick breakdown: - The four inputs are sent to LLMs with precise instructions to generate: first, prompts for images and image modifications; second, prompts for animations; third, prompts for generating music. - All voices are generated from the text and timed precisely, as they determine the length of each animation segment. - The first image and video are generated to serve as the title, but also as the guide for all other images created for the video. - Titles and subtitles are also added automatically in Comfy. - I also developed a lot of custom nodes for minor frame calculations, mostly to match audio and video. - The full system is a large loop that, for each line of text, generates an image and then a video from that image. The loop was the hardest part to build in this workflow, so it can process either a 20-second video or a 2-minute video with the same input. - There are multiple combinations of LLMs that try to understand the text in the best way to provide the best prompts for images and video. - The final video is assembled entirely within ComfyUI. - The music is generated based on the LLM output and matches the exact timing of the full animation. - Done! For reference, this workflow uses a lot of models and only works on an RTX 6000 Pro with plenty of RAM. My goal is not to replace humans, as I’ll try to explain later, this workflow is highly controlled and can be adapted or reworked at any point by real artists! My aim was to create a tool that can animate text in one go, allowing the AI some freedom while keeping a strict flow. I don’t know yet how I’ll share this workflow with people, I still need to polish it properly, but maybe through Patreon. Anyway, I hope you enjoy my research, and let’s always keep pushing further! :)

Lovis Odin

58,841 görüntüleme • 11 ay önce

OpenAI released Sora 2, their new state of the art video and audio generation model, calling it the GPT-3.5 moment for video - The model generates videos up to 10 seconds long (default 9:16 vertical) with synchronized audio including difficult physics simulations, where mistakes the model makes frequently appear to be mistakes of the internal agent that Sora 2 is implicitly modeling rather than physics-breaking errors - Cameos let users record a short one-time video-and-audio capture where they read a verification phrase, then the model can insert them into any Sora-generated environment with accurate appearance and voice, where only they decide who can use their cameo and they can revoke access or remove any video that includes it at any time including drafts created by other people - The app functions as a social iOS platform not optimized for time spent but explicitly designed to maximize creation not consumption, heavily favoring content from people you follow or interact with and videos the model thinks will inspire your own creations, with features like remixing other posts and direct messaging to share videos privately - Safety measures include limited invitation rollout, restricting image uploads that feature a photorealistic person, blocking all video uploads, no video-to-video transformation at launch, stricter safeguards for minors with default scroll limits, C2PA metadata and visible watermarks on all outputs, and evaluations showing 95.1-99.7% effectiveness at blocking unsafe content - Initial rollout covers the US and Canada with plans to expand while the UK, EU, and Australia are not included at launch, available for free with generous limits though subject to compute constraints with the only planned monetization being optional payment for extra videos if demand exceeds compute, ChatGPT Pro users get access to higher quality Sora 2 Pro model and API release planned for future

Tibor Blaho

62,482 görüntüleme • 10 ay önce

🎥 Introducing Hailuo's Subject Reference: Revolutionizing Character Consistency in Video Creation 🔥 We’re excited to present Hailuo's S2V-01 model, a groundbreaking innovation in AI video generation that tackles one of the industry’s biggest challenges: maintaining consistent, realistic facial features and identity across dynamic video content, regardless of camera angles or movements. 💡 Why It’s a Game Changer: - Pioneering Technology: The first-of-its-kind to ensure character consistency in dynamic video generation, surpassing even fine-tuned models in performance. - Minimal Input, Maximum Impact: Generate character-consistent videos from just one reference image. Every frame remains true to the original identity with unmatched accuracy and reliability. - Enhanced Flexibility: Adjust more than just facial features—modify posture, expressions, lighting, and more, all with simple text-based prompts. 🌟While the new model enhances subject consistency, it may occasionally follow prompts less precisely than T2V or I2V, with some environmental morphing. Despite these early-stage challenges, Hailuo Subject Reference marks a significant leap in AI video generation. We’re committed to continual improvements including multi-subject references, objects references, and complex, multi-layered scenes. Explore the future of creative, consistent video production with Hailuo S2V-01 today. 🔥We believe the possibilities are endless.

Hailuo AI-MiniMax Hub

692,515 görüntüleme • 1 yıl önce

Elon Musk just made one if the biggest moves in taking over the programming industry “SpaceX just bought Cursor for $60 billion. Do you realize how big this is? SpaceX went public — the biggest IPO in history. $75 billion raised, almost a $2 trillion valuation and the first thing to do with that money? Buy the most popular AI coding tool on the planet. Here's why that changes everything. Elon now owns 3 layers: the compute, Colossus data centers, the models, Grok through xAI, and now the tool that developers actually use every day. It's the full stack. And here's what makes Cursor different from Claude Code or Codex. Cursor is model agnostic. You can run Claude in it, GPT, Gemini, whatever model you want. It's not locked to any one company, and now it has SpaceX's resources behind it. Cursor said they were bottlenecked by compute. Well, that bottleneck has just been removed. $4 billion in annual revenue, over half the Fortune 500 already uses it, and now it's backed by a $2 trillion company. OpenAI has Codex, Anthropic has Claude Code, and now Elon has Cursor.” Let me break this down in simple terms Elon Musk now controls more of the full AI picture: - Massive computers, power (data centers like Colossus) - Smart AI models (Grok from xAI) - The actual tool millions of developers use every day (Cursor) For every day users this means Faster and smarter apps and websites in the future. More developers using powerful AI tools means new apps, games, websites, and features get built quicker and cheaper. This means better video games, smoother streaming, smarter phone apps and better programs For Developers they can describe what they want in plain English (“make a feature that does X”) and the AI handles more of the heavy lifting

Wall Street Apes

213,589 görüntüleme • 2 ay önce