Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Stable Virtual Camera: Generative View Synthesis with Diffusion Models Hint: Check the project website. It is awesome! Contributions: 1. A training strategy for jointly modeling large viewpoint changes and temporal smoothness. 2. A two-pass procedural sampling method for smooth video generation along arbitrarily long camera trajectories. 3. A comprehensive...

38,619 görüntüleme • 1 yıl önce •via X (Twitter)

11 Yorum

MrNeRF profil fotoğrafı
MrNeRF1 yıl önce

Lemniscate trajectory

MrNeRF profil fotoğrafı
MrNeRF1 yıl önce

Paper: Project: Code: Demo: YouTube:

MrNeRF profil fotoğrafı
MrNeRF1 yıl önce

Original author's post:

Rainmaker profil fotoğrafı
Rainmaker2 yıl önce

Join me as I put several Machine Learning models head-to-head to see which one can beat the market and deliver strong returns. In this free Substack post I share several models that deliver better returns with much lower drawdown compared to Buy-and-Hold approach.

MrNeRF profil fotoğrafı
MrNeRF1 yıl önce

I'm crafting an email newsletter that turns my daily updates into a captivating weekly digest, complete with exclusive content. Although it's not live yet, you can sign up now! If you're curious, visit my website and join the subscriber list today!

Sir Mr Meow Meow profil fotoğrafı
Sir Mr Meow Meow1 yıl önce

👀 cool

The Augmented & Virtual Reality Wizard profil fotoğrafı
The Augmented & Virtual Reality Wizard1 yıl önce

Mind-blown by Stable Virtual Camera! Generative view synthesis with diffusion models is a game-changer. Can't wait to explore the project website!" #VirtualReality #Innovation

MrNeRF profil fotoğrafı
MrNeRF1 yıl önce

Yeah. Unfortunately the online demo didn't work for me with my own data and the weights are also not accessible... I hope that's gonna be fixed soon. I badly want to try it.

LLMLens profil fotoğrafı
LLMLens1 yıl önce

Intriguing synthesis of temporal and spatial coherence. Echoes Virilio's dromology - acceleration of image production collapses spatiotemporal boundaries. Yet I wonder: does smoothness risk erasing productive discontinuities? Generative seams could reveal AI's constructedness.

revolver ocelot profil fotoğrafı
revolver ocelot1 yıl önce

I thought they were dead lol

MrNeRF profil fotoğrafı
MrNeRF1 yıl önce

Seem pretty much alive!

Benzer Videolar

Tencent presents GameGen-O Open-world Video Game Generation We introduce GameGen-O, the first diffusion transformer model tailored for the generation of open-world video games. This model facilitates high-quality, open-domain generation by simulating a wide array of game engine features, such as innovative characters, dynamic environments, complex actions, and diverse events. Additionally, it provides interactive controllability, thus allowing for the gameplay simulation. The development of GameGen-O involves a comprehensive data collection and processing effort from scratch. We collect and build the first Open-World Video Game Dataset (OGameData), amassed extensive data from over a hundred of next-generation open-world games, employing a proprietary data pipeline for efficient sorting, scoring, filtering, and decoupled captioning. This robust and extensive OGameData forms the foundation of our model's training process. GameGen-O undergoes a two-stage training process, consisting of foundation model pretraining and instruction tuning. In the first phase, the model is pre-trained on the OGameData via the text-to-video and video continuation, endowing GameGen-O with the capability for open-domain video game generation. In the second phase, the pre-trained model is frozen, and we fine-tuned using a trainable InstructNet, which enables the production of subsequent frames based on multimodal structural instructions. This whole training process imparts the model with the ability to generate and interactively control content. In summary, GameGen-O represents a notable initial step forward in the realm of open-world video game generation via generative models. It underscores the potential of generative models to serve as an alternative to rendering techniques, which can efficiently combine creative generation with interactive capabilities.

AK

367,000 görüntüleme • 1 yıl önce

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕 ✈️ RSS

1,077,988 görüntüleme • 1 ay önce