Loading video...

Video Failed to Load

Go Home

Today we released the code for our CVPR 2026 paper, Flowception. Flowception bridges fully bidirectional sequence modeling and autoregressive generation by inserting frames via learned order, then denoising them with continuous flow. Website: Code:

18,980 views • 2 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Depth Any Video with Scalable Synthetic Data AI physicists and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.

MrNeRF

27,428 views • 1 year ago

Up, Up, Down, Down, Left, Right, Left Right, B, A…❤️ As a kid I called the Konami Code the “Contra Code.” I played with friends and instinctively recall “Select, Start” at the end, which gives 30 lives for 2 players. Yet “Select” scrolls to the 2 player option and “Start” is necessary to start the game. Thus, it’s technically not part of the code. Pressing only “Start” after goes to 1 player mode with 30 lives. Notably, the 1st Nintendo Power issue’s “Classified Information” section explicitly included “Start” when disclosing the sequence. The Konami Code was initially created by Kazuhisa Hashimoto while testing the arcade to console port of Gradius, which he found too difficult. (It also “includes” Start to start the game.) There, it gave access to all power ups ups. Why it stayed in Gradius is unclear, and it appears in future Konami games, most famously Contra. It confers 30 lives in Contra, Life Force, and others, and is sometimes called “the 30 Lives Code.” The effect can vary in different games. As kids, knowing the code was like magic. We imparted it to our friends like a fabled wisdom. Like riding a bike, I’ll never forget it, and I smile inputting it now just as much as I did as a child. No shame in it—Nintendo codes were and still are rad. Homages to the Konami Code outside of gaming are ubiquitous, including in Wreck it Ralph, unlocking Amazon’s “Super Alexa,” and even unlocking a chiptune version of the National Anthem on a Bank of Canada website! Also, Contra’s soundtrack is so good…I hope I’m not the only dork who takes a moment to jam out to it 🤭

Christina Rose

14,993 views • 1 year ago

Introducing Kaleido💮 from AI at Meta — a universal generative neural rendering engine for photorealistic, unified object and scene view synthesis. Kaleido is built on a simple but powerful design philosophy: 3D perception is a form of visual common sense. Following this idea, we formulate rendering purely as a sequence-to-sequence generation problem, successfully unifying neural rendering with the architecture principles behind modern language and video models. Unlike traditional neural rendering methods, Kaleido learns 3D purely in a data-driven way, without explicit 3D representations or structures. It acquires spatial understanding directly through large-scale video pretraining, then multi-view 3D data finetuning, inspired by how LLMs acquire textual common sense from large corpora before specialising in domains like coding. Through extensive ablations, we progressively modernised the architecture design and training strategies and tackled key scaling challenges in sequence-to-sequence generative rendering, arriving at a design that’s simple, versatile, and scalable. Kaleido significantly outperforms prior generative models in few-view settings, and remarkably is the first zero-shot generative method matches InstantNGP-level rendering quality in multi-view settings. We view Kaleido also as an alternative step towards world modeling that flexibly spans a spectrum of “realities": with many views, it faithfully reconstructs grounded reality; with fewer views, it imagines plausible unseen details. 🔗 Explore more results and paper:

Shikun Liu

22,332 views • 9 months ago