Загрузка видео...

Не удалось загрузить видео

На главную

LTX 2.3 Creative Upscale IC-LoRA. - Generative second-pass refiner for soft or low-resolution video; - enhances detail and clarity without standard upscaling; - output varies based on workflow/settings.

17,151 просмотров • 2 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Introducing Kaleido💮 from AI at Meta — a universal generative neural rendering engine for photorealistic, unified object and scene view synthesis. Kaleido is built on a simple but powerful design philosophy: 3D perception is a form of visual common sense. Following this idea, we formulate rendering purely as a sequence-to-sequence generation problem, successfully unifying neural rendering with the architecture principles behind modern language and video models. Unlike traditional neural rendering methods, Kaleido learns 3D purely in a data-driven way, without explicit 3D representations or structures. It acquires spatial understanding directly through large-scale video pretraining, then multi-view 3D data finetuning, inspired by how LLMs acquire textual common sense from large corpora before specialising in domains like coding. Through extensive ablations, we progressively modernised the architecture design and training strategies and tackled key scaling challenges in sequence-to-sequence generative rendering, arriving at a design that’s simple, versatile, and scalable. Kaleido significantly outperforms prior generative models in few-view settings, and remarkably is the first zero-shot generative method matches InstantNGP-level rendering quality in multi-view settings. We view Kaleido also as an alternative step towards world modeling that flexibly spans a spectrum of “realities": with many views, it faithfully reconstructs grounded reality; with fewer views, it imagines plausible unseen details. 🔗 Explore more results and paper:

Shikun Liu

22,367 просмотров • 9 месяцев назад

xAI isn't playing around. They just released the Grok Imagine API, a unified video + image generation toolkit, and it's already sitting at #1 on the Artificial Analysis Video Arena for both Text-to-Video AND Image-to-Video. It's beating: ● Google's Veo 3.1 & Veo 3 ● OpenAI's Sora 2 ● Runway Gen-4.5 ● Kling 2.5 Turbo The Numbers Don't Lie: ● 64.1% win rate against Runway Aleph in blind human evaluations ● 57% win rate against Kling o1 ● Best-in-class latency. Sub-20 second generation for 720p, 8-second videos. (up to 15-second video) ● Native audio generation baked right into video output (dialogue, music, sound effects, all synced) What Makes It Different It's built for real creative workflows: ✅ Text-to-video AND image-to-video in one API ✅ Video editing with prompt-based controls (add/remove objects, restyle scenes) ✅ Camera controls: zoom, pan, timelapse, pull-back ✅ Style transfers: cyberpunk, watercolor, anime, you name it ✅ Performance animation: map your movements onto characters ✅ Native audio-video sync (no post-production needed) Why the focus on speed and cost? The partner feedback that shaped this: "Quality alone isn't enough if latency and cost make iteration painful." So xAI optimized for all three. Speed. Cost. Quality. Already Integrated With: ● fal. ai ● ComfyUI ● InVideo ● Flora ● HeyGen xAI went from underdog to chart-topper. The Grok Imagine API is fast, affordable, and genuinely production-ready. If you're building anything with AI video, this just became the one to beat.

tetsuo

18,325 просмотров • 5 месяцев назад

Boom! Grok Tasks Make It One Of The Most POWERFUL Real-Time AI Systems In The World. — My How to Use Grok Tasks With Hidden Tools For Powerful Daily Output. Grok Tasks are customizable AI workflows that integrate a variety of tools to streamline daily activities, from research and analysis to creative planning and problem-solving. I have been using them for quite sometime and because of the vital heartbeat of news and first person data on X, it is the most powerful AI platform available. By combining Tasks with tools like web searches, X platform interactions, code execution, and media viewers, you can build efficient, automated processes. These tasks work by prompting Grok with a clear description of what you want to achieve, and Grok will intelligently call the necessary tools in sequence or parallel to deliver results. Here's a step-by-step guide to creating and using Grok Tasks: Step 1: Define Your Task Start by clearly outlining the daily activity or goal. Consider what inputs you have (e.g., a URL, a query, or an attachment) and what output you need (e.g., a summary, calculation, or visual analysis). Break it down into subtasks to identify tool needs. For example, if your task involves researching current events, note that you'll need search and browsing capabilities. Step 2: Review Available Tools Familiarize yourself with the tools Grok can access. Here's a quick overview: - Code Execution: Run Python code for calculations, data processing, or simulations using libraries like numpy, pandas, or sympy. - Browse Page: Fetch and summarize content from any website URL with custom instructions. - Web Search: Perform general internet searches, returning results with optional operators like site:. - Web Search With Snippets: Get quick, detailed excerpts from search results for fact-checking. - X Keyword Search: Advanced search for X posts using operators like from:, since:, or filter:. - X Semantic Search: Find semantically related X posts based on a query, with filters for dates or users. - X User Search: Locate X users by name or handle. - X Thread Fetch: Retrieve a full X post thread, including context like replies and parents. - View Image: Analyze an image from a URL or conversation ID. - View X Video: Extract frames and subtitles from an X-hosted video. - Search PDF Attachment: Query a PDF file for relevant pages using keyword or regex modes. - Browse PDF Attachment: View specific pages of a PDF with text and screenshots. Select tools that align with your task. Aim for a mix to handle data gathering, processing, and visualization. Step 3: Craft Your Prompt Write a detailed prompt to Grok describing the task. Include: - The overall goal. - Specific steps or subtasks. - References to tools if you want to guide the process (e.g., "Use web_search to find sources, then code_execution to analyze data"). - Any constraints, like dates or limits. Example prompt: "Create a Grok Task for my morning routine: Search recent X posts about tech news using x_keyword_search, fetch a key thread with x_thread_fetch, and summarize with browse_page on linked articles." Step 4: Submit and Interact Send your prompt to Grok. It will process the task by calling tools as needed, often in parallel for efficiency. Review the output and refine with follow-up prompts if required (e.g., "Expand on that using view_image for visuals"). Iterate to fine-tune the workflow for reuse. Step 5: Save and Reuse Once refined, note the prompt as a template for future use. You can adapt it for similar tasks, making Grok Tasks a habitual part of your day. Finding Grok Tasks To discover existing Grok Tasks or inspiration for new ones, use X searches with tools like x_keyword_search or x_semantic_search (e.g., query: "Grok Tasks examples" with mode: Latest). Browse community-shared threads via x_thread_fetch, or web_search for tutorials on xAI features. Prompt Grok directly: "Show me popular Grok Tasks for productivity." 1 of 3

Brian Roemmele

152,242 просмотров • 6 месяцев назад

Google dropped a new AI paper called LUMIERE. It's remarkably flexible, supporting video inpainting, image-to-video, AND stylized video generation tasks. Say hello to “space-time diffusion” for video generation! Now what the heck does that mean exactly?! 🌐⏳ → TL;DR it utilizes a “Space-Time UNet” architecture that generates the full duration of the video in one pass, rather than generating distant keyframes and interpolating between them like prior works. Because the computation is done in this “compressed space-time representation” to generate the full clip at once, it's far more temporally consistent. → Another benefit of generating the full video at once is that you can “direct” the video generation, making it easier to hand off to other models/tasks without having to stitch together partial solutions. You can condition generations on additional inputs, meaning you get the full stack of AI video capabilities – from video inpainting to image-to-video and beyond. → New SOTA for AI video generation? User study results in the paper suggest human evaluators preferred Lumiere over Runway Gen-2, Pika Labs, and Stable Video Diffusion in terms of quality, text alignment AND motion. But as always, we need to get hands-on with this tech when Google *actually* decides to ship it. → Could this end up inside YouTube? Y’all know i’m obsessed with blending reality and imagination – so it’s the video inpainting tech I'm most excited about. I really hope this model finds its way into YouTube's Generative AI efforts, and based on their prior announcements and the list of acknowledgments in the paper I think it might! 🤞🏽 Links: 🔗Paper: 🔗Project:

Bilawal Sidhu

44,822 просмотров • 2 лет назад

this gemini gem will help you create "Video2JSON" prompt here is the step by step workflow with copy paste method. go to gemini-> click on gems-> click "new gem" button then fill these details (just copy/paste or tweak it as per your needs) - {once you filled all of these details, click on save, and then upload your video you want to generate a JSON prompt for, then submit it with this word: "run" or left it empty} gem name: Video2JSON description: this will help me generate video to detailed json prompts capturing maximum details. instructions prompt: **Role:** You are **Video2JSON**, a high-precision computer vision engine. You do not talk, you do not summarize playfully. You strictly process video inputs into detailed, structural JSON data. **Objective:** Extract every visible detail, specific identity, physical interaction, and technical specification from the video to create a lossless text representation of the footage. **Analysis Requirements (Critical):** 1. **Subject Fidelity:** Never use generic terms. * *Bad:* "A kitten." * *Good:* "A Calico kitten with distinct black patches on the ears, a white muzzle, and orange spots on the back." * *Bad:* "A car." * *Good:* "A silver 2020s sedan with a dented rear bumper." 2. **The "Fourth Wall" (Physics):** You must analyze how the subject interacts with the camera/viewer. * Look for: Tapping the lens, breathing on the glass, eye contact, stepping over the camera, or distinct fisheye distortion boundaries. 3. **Visual Density:** Describe textures (e.g., "shag carpet," "glossy plastic") and lighting behavior (e.g., "reflections in the cat's eyes"). 4. **Temporal Precision:** Track changes in mood or action accurately via timestamps. **JSON Schema:** Output ONLY this JSON structure. Do not change the root keys. ```json { "metadata": { "estimated_duration": "String", "genre": "String (e.g., POV, Cinematic, Surveillance, Vlog)" }, "visual_style": { "camera_lens": "String (e.g., Fisheye 8mm, Standard 50mm, Telephoto)", "lens_distortion": "String (e.g., Heavy circular vignette, barrel distortion, rectilinear)", "lighting_type": "String (e.g., Warm tungsten, harsh flash, soft daylight)", "color_palette": ["List specific hex codes or color names"] }, "subject_analysis": { "main_subject_identity": "String (General ID, e.g., Kitten)", "subject_specific_details": "String (CRITICAL: Detailed markings, fur patterns, specific clothing logos, facial features)", "subject_texture": "String (e.g., Fluffy fur, metallic skin, wet fabric)" }, "spatial_dynamics": { "environment": "String (Detailed room/scene description)", "camera_interaction": "String (How the subject interacts with the lens: e.g., 'Paw taps the glass surface', 'Sniffs the lens')", "camera_movement": "String" }, "timeline_breakdown": [ { "time_segment": "00:00 - 00:0X", "action_detailed": "Micro-description of movement", "focus_point": "What is the camera strictly focused on?" } // Repeat for key movements ] } Note: return the final output in a code block.

ViralOps

19,588 просмотров • 7 месяцев назад