Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Stable Virtual Camera: Generative View Synthesis with Diffusion Models Hint: Check the project website. It is awesome! Contributions: 1. A training strategy for jointly modeling large viewpoint changes and temporal smoothness. 2. A two-pass procedural sampling method for smooth video generation along arbitrarily long camera trajectories. 3. A comprehensive...

38,619 Aufrufe • vor 1 Jahr •via X (Twitter)

11 Kommentare

Profilbild von MrNeRF
MrNeRFvor 1 Jahr

Lemniscate trajectory

Profilbild von MrNeRF
MrNeRFvor 1 Jahr

Paper: Project: Code: Demo: YouTube:

Profilbild von MrNeRF
MrNeRFvor 1 Jahr

Original author's post:

Profilbild von Rainmaker
Rainmakervor 2 Jahren

Join me as I put several Machine Learning models head-to-head to see which one can beat the market and deliver strong returns. In this free Substack post I share several models that deliver better returns with much lower drawdown compared to Buy-and-Hold approach.

Profilbild von MrNeRF
MrNeRFvor 1 Jahr

I'm crafting an email newsletter that turns my daily updates into a captivating weekly digest, complete with exclusive content. Although it's not live yet, you can sign up now! If you're curious, visit my website and join the subscriber list today!

Profilbild von Sir Mr Meow Meow
Sir Mr Meow Meowvor 1 Jahr

👀 cool

Profilbild von The Augmented & Virtual Reality Wizard
The Augmented & Virtual Reality Wizardvor 1 Jahr

Mind-blown by Stable Virtual Camera! Generative view synthesis with diffusion models is a game-changer. Can't wait to explore the project website!" #VirtualReality #Innovation

Profilbild von MrNeRF
MrNeRFvor 1 Jahr

Yeah. Unfortunately the online demo didn't work for me with my own data and the weights are also not accessible... I hope that's gonna be fixed soon. I badly want to try it.

Profilbild von LLMLens
LLMLensvor 1 Jahr

Intriguing synthesis of temporal and spatial coherence. Echoes Virilio's dromology - acceleration of image production collapses spatiotemporal boundaries. Yet I wonder: does smoothness risk erasing productive discontinuities? Generative seams could reveal AI's constructedness.

Profilbild von revolver ocelot
revolver ocelotvor 1 Jahr

I thought they were dead lol

Profilbild von MrNeRF
MrNeRFvor 1 Jahr

Seem pretty much alive!

Ähnliche Videos

Tencent presents GameGen-O Open-world Video Game Generation We introduce GameGen-O, the first diffusion transformer model tailored for the generation of open-world video games. This model facilitates high-quality, open-domain generation by simulating a wide array of game engine features, such as innovative characters, dynamic environments, complex actions, and diverse events. Additionally, it provides interactive controllability, thus allowing for the gameplay simulation. The development of GameGen-O involves a comprehensive data collection and processing effort from scratch. We collect and build the first Open-World Video Game Dataset (OGameData), amassed extensive data from over a hundred of next-generation open-world games, employing a proprietary data pipeline for efficient sorting, scoring, filtering, and decoupled captioning. This robust and extensive OGameData forms the foundation of our model's training process. GameGen-O undergoes a two-stage training process, consisting of foundation model pretraining and instruction tuning. In the first phase, the model is pre-trained on the OGameData via the text-to-video and video continuation, endowing GameGen-O with the capability for open-domain video game generation. In the second phase, the pre-trained model is frozen, and we fine-tuned using a trainable InstructNet, which enables the production of subsequent frames based on multimodal structural instructions. This whole training process imparts the model with the ability to generate and interactively control content. In summary, GameGen-O represents a notable initial step forward in the realm of open-world video game generation via generative models. It underscores the potential of generative models to serve as an alternative to rendering techniques, which can efficiently combine creative generation with interactive capabilities.

AK

367,000 Aufrufe • vor 1 Jahr

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕

1,078,049 Aufrufe • vor 1 Monat