Загрузка видео...

Не удалось загрузить видео

На главную

LTX 2.3 Creative Upscale IC-LoRA. - Generative second-pass refiner for soft or low-resolution video; - enhances detail and clarity without standard upscaling; - output varies based on workflow/settings.

17,151 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Introducing Kaleido💮 from AI at Meta — a universal generative neural rendering engine for photorealistic, unified object and scene view synthesis. Kaleido is built on a simple but powerful design philosophy: 3D perception is a form of visual common sense. Following this idea, we formulate rendering purely as a sequence-to-sequence generation problem, successfully unifying neural rendering with the architecture principles behind modern language and video models. Unlike traditional neural rendering methods, Kaleido learns 3D purely in a data-driven way, without explicit 3D representations or structures. It acquires spatial understanding directly through large-scale video pretraining, then multi-view 3D data finetuning, inspired by how LLMs acquire textual common sense from large corpora before specialising in domains like coding. Through extensive ablations, we progressively modernised the architecture design and training strategies and tackled key scaling challenges in sequence-to-sequence generative rendering, arriving at a design that’s simple, versatile, and scalable. Kaleido significantly outperforms prior generative models in few-view settings, and remarkably is the first zero-shot generative method matches InstantNGP-level rendering quality in multi-view settings. We view Kaleido also as an alternative step towards world modeling that flexibly spans a spectrum of “realities": with many views, it faithfully reconstructs grounded reality; with fewer views, it imagines plausible unseen details. 🔗 Explore more results and paper:

Shikun Liu

22,471 просмотров • 11 месяцев назад

xAI isn't playing around. They just released the Grok Imagine API, a unified video + image generation toolkit, and it's already sitting at #1 on the Artificial Analysis Video Arena for both Text-to-Video AND Image-to-Video. It's beating: ● Google's Veo 3.1 & Veo 3 ● OpenAI's Sora 2 ● Runway Gen-4.5 ● Kling 2.5 Turbo The Numbers Don't Lie: ● 64.1% win rate against Runway Aleph in blind human evaluations ● 57% win rate against Kling o1 ● Best-in-class latency. Sub-20 second generation for 720p, 8-second videos. (up to 15-second video) ● Native audio generation baked right into video output (dialogue, music, sound effects, all synced) What Makes It Different It's built for real creative workflows: ✅ Text-to-video AND image-to-video in one API ✅ Video editing with prompt-based controls (add/remove objects, restyle scenes) ✅ Camera controls: zoom, pan, timelapse, pull-back ✅ Style transfers: cyberpunk, watercolor, anime, you name it ✅ Performance animation: map your movements onto characters ✅ Native audio-video sync (no post-production needed) Why the focus on speed and cost? The partner feedback that shaped this: "Quality alone isn't enough if latency and cost make iteration painful." So xAI optimized for all three. Speed. Cost. Quality. Already Integrated With: ● fal. ai ● ComfyUI ● InVideo ● Flora ● HeyGen xAI went from underdog to chart-topper. The Grok Imagine API is fast, affordable, and genuinely production-ready. If you're building anything with AI video, this just became the one to beat.

tetsuo

18,325 просмотров • 7 месяцев назад

Shipped. Minimax H3 EZLaunch. One install script. Full optimized stack for RTX 3090 and RTX 4090. Windows AND Linux. Ships with every model you need. T2V, I2V, and Ref2V all supported. 5 second clips in about 90 seconds. 15 second clips in under 8 minutes. Zero OOMs across 80 minutes of continuous generation. Same stack that held clean across all 10 space battle renders. What it actually does: - Downloads and configures the correct weight set automatically: pruned INT8 UNET with ConvRot, Turbo LoRA (8-step quality band), Qwen3-VL text encoder in NVFP4, MiniMax video + audio VAEs. No hunting through HuggingFace for the right quants. - Full mode coverage. Text-to-video out of the box. Image-to-video with first/last frame. Reference-to-video for style-guided generation. O - Patches SageAttention v2 sm89 to Triton INT8. - Enables kitchen CUDA with FORCE_CUDA on torch cu128 + driver 580 - Full flag set applied: lowvram, disable-smart-memory, disable-pinned-memory, disable-cuda-malloc, fp16-intermediates, CLIP on CPU - Timed smoke test runner included. Verify your stack without guessing. - Cross-platform installer. Same script, same results, Windows or Linux. The weight selection alone saves people hours. The Sage patch is the difference between launch-fail and sub-90 second renders on Ada. Nobody had either documented publicly. Now it is one command. Try it on similar cards and tweak if you want. MVP. More GPU configurations coming. Stay Tuned. Repo link in reply.

Yume_X

13,279 просмотров • 1 месяц назад