Most video models 🤯forget the past 🐌slow down over... time 🔁rely on bidirectional (not causal) attention Our state-space video world models (SSM) 🧠remember across hundreds of frames ⚡️generate at constant speed ⏩is fully causal, enabling real-time rollout 1/3show more

Gordon Wetzstein
20,043 次观看 • 1 年前
(1/n) 🚀 With FastVideo, you can now generate a... 5-second video in 5 seconds on a single H200 GPU! Introducing FastWan series, a family of fast video generation models trained via a new recipe we term as “sparse distillation”, to speed up video denoising time by 70X! 🖥️ Live demo: (Thanks to @gmicloud for the support!) 🔗 Blog: 🔓 We fully open-source our models, code, and data with Apache-2.0 licensesshow more

Hao AI Lab
78,660 次观看 • 1 年前
NVIDIA just released a very impressive text-to-video paper. Video... Latent Diffusion Models (Video LDMs) use a diffusion model in a compressed latent space to generate high-resolution videos. Here's a brief overview of how it works: 1. Pre-train image LDM on a dataset of images. 2. Turn the image LDM into a Video LDM by adding temporal layers to model video frames. 3. Fine-tune the Video LDM on encoded video sequences to create a video generator. 4. Temporally align diffusion model upsamplers to generate high-resolution videos. 5. Validate Video LDM on real driving videos of 512x1024 resolution, achieving state-of-the-art performance. 6. Apply the approach in creative content creation with text-to-video modeling. Paper: Project:show more

Lior Alexander
158,595 次观看 • 3 年前
AI Is Moving Beyond “Generating Videos” — Toward “Generating... Worlds” Over the past two years, AI video models have advanced at an astonishing pace. From Runway and Pika to Sora and Veo, AI-generated videos have become increasingly realistic and more consistent with the physical laws of the real world. Many people believe the next objective is simply to generate videos that are longer, sharper, and more lifelike. But if we take a step back, we can see that the real transformation is not happening in video itself. It is happening in world models. What Is a World Model? In 1943, psychologist Kenneth Craik proposed an idea that would influence artificial intelligence research for decades. He argued that the human brain does not merely react to the outside world. Instead, it maintains an internal model of how the world works. Because we have this internal model, we can predict the outcome of an action before we actually take it. Before crossing a road, we estimate whether a car will pass by. Before catching a ball, we predict its trajectory. These abilities come from continuously simulating the world in our minds, rather than relying entirely on trial and error. This idea later became known by a more formal term: World Model. A world model does not describe a single image or a fixed video clip. It is an internal representation capable of continuously simulating the rules and dynamics of the real world. Why Is AI Research Turning Toward World Models? Because predicting “what comes next” is becoming increasingly central to how AI systems work. Language models predict the next token. Image models predict the next step in the denoising process. Video models predict the next frame. A world model, however, attempts to predict something broader: What should the world look like in the next moment? In 2018, David Ha and Jürgen Schmidhuber proposed in their paper World Models that an intelligent agent could first learn a model of the world, and then use that internal model to plan its actions. The Dreamer series later demonstrated that many complex tasks could be learned by training agents inside an “imagined world.” At the same time, the development of video models such as Sora and Veo led researchers to another realization: A model capable of continuously generating video has already learned, at least implicitly, many of the rules governing the real world. As a result, these two research directions have gradually begun to converge. But Video Is Not Yet a World This is where the distinction is often misunderstood. For a world model to support meaningful real-time interaction, it must solve several critical problems. Most video models today are essentially answering one question: What should the next frame look like? A true world model needs to answer much more: What happens if I take one step forward? If I walk behind a building and then return, will the building still be there? If I suddenly change the camera angle, will the entire space remain consistent? If I enter a command such as: “Summon a dragon.” Will the world respond immediately? In other words, a world model must do more than generate content. It must understand space. It must understand time. It must understand causality. And it must understand interaction. Moving from watching to participating is where the real difficulty of world models begins. World Models Are Entering the Interactive Era One of the latest attempts in this direction is Alaya World, recently open-sourced by Alaya World, or Alaya Lab. Instead of generating a fixed video clip, it generates a world that users can explore in real time. Users can begin with text, an image, or a video, enter the generated scene, move freely through it, and introduce new prompts at any moment during generation. The world responds immediately. According to the publicly released information, Alaya World provides: Real-time streaming generation at 720p and 24 FPS Stable continuous exploration for more than one minute The ability to switch prompts and trigger skills or events during generation Model weights and inference code released under the Apache 2.0 License Training code and datasets planned for future release What makes these capabilities important is not simply the technical specifications. It is that the generated “world” can now support continuous interaction. The official demo shows that users can genuinely control, transform, and explore the generated environment. AI Is Evolving From a Tool Into an Environment Over the past few years, most discussions around AI have focused on content generation. Generating text. Generating images. Generating videos. But world models raise a fundamentally different question: Can AI generate an environment that people can inhabit, explore, and continuously evolve? If the answer is yes, the impact will extend far beyond video generation. Game development, robotics training, embodied intelligence, digital twins, virtual production, and many other fields could be transformed by the development of world models. World models are still at a very early stage. Yet from Craik’s proposal of an internal mental model more than eighty years ago to the emergence of today’s interactive world-generation systems, a clear evolutionary path is beginning to take shape. Perhaps what AI is ultimately learning has never been limited to images, videos, or language. Perhaps it is learning the world itself. References GitHub: Technical Report:show more

雪踏乌云
113,347 次观看 • 1 个月前
Depth Any Video with Scalable Synthetic Data AI physicists... and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.show more

MrNeRF
27,428 次观看 • 1 年前
Every Mega Hero is a personalized NFT, paired with... a live trading bot fully configured by the player. The minting process unfolds in 4 steps: 1 - Spirit defines risk appetite: • Zen (low risk) • Dynamic (balanced) • Degen (max chaos) 2 - Weapons : up to 5 trading pairs selected (trading arsenal) 3 - Strategies : up to five strategies chosen from a curated pool of quant models, optimized for various market conditions 4 - Sparks (behavioral traits): • Resilience (hedging) • Vigilance (reaction speed) • Time Travel (historical pattern analysis) Once minted, the Hero runs a $10K virtual portfolio in real-time. At any point, it can be linked to a real trading account mirroring its virtual trades in the real world. A new class of trading companion is coming. Mega Heroes landing soon on MegaETH.show more

MegaHeroes | ∑:
11,361 次观看 • 1 年前
Most video tools can generate clips. Very few can... maintain identity. That has been the real bottleneck in AI video creation. Kling O1 changes that. For the first time, creators can carry a character, style, and visual language across scenes without constant fixes. You can reference past clips, assets, or images and the output stays consistently on-model. No visual drift. No rework loops. No “this doesn’t look like the last shot” moments. It feels less like prompting a tool and more like working with a creative collaborator that remembers context. The impact is practical, not theoretical: → Faster production cycles → Lower iteration costs → Noticeably higher output quality This is what mature AI tooling looks like. Not louder features. Not bigger claims. Just reliability where it actually matters. Consistency is no longer the problem.show more

Darshal Jaitwar
141,038 次观看 • 7 个月前
Google dropped a new AI paper called LUMIERE. It's... remarkably flexible, supporting video inpainting, image-to-video, AND stylized video generation tasks. Say hello to “space-time diffusion” for video generation! Now what the heck does that mean exactly?! 🌐⏳ → TL;DR it utilizes a “Space-Time UNet” architecture that generates the full duration of the video in one pass, rather than generating distant keyframes and interpolating between them like prior works. Because the computation is done in this “compressed space-time representation” to generate the full clip at once, it's far more temporally consistent. → Another benefit of generating the full video at once is that you can “direct” the video generation, making it easier to hand off to other models/tasks without having to stitch together partial solutions. You can condition generations on additional inputs, meaning you get the full stack of AI video capabilities – from video inpainting to image-to-video and beyond. → New SOTA for AI video generation? User study results in the paper suggest human evaluators preferred Lumiere over Runway Gen-2, Pika Labs, and Stable Video Diffusion in terms of quality, text alignment AND motion. But as always, we need to get hands-on with this tech when Google *actually* decides to ship it. → Could this end up inside YouTube? Y’all know i’m obsessed with blending reality and imagination – so it’s the video inpainting tech I'm most excited about. I really hope this model finds its way into YouTube's Generative AI efforts, and based on their prior announcements and the list of acknowledgments in the paper I think it might! 🤞🏽 Links: 🔗Paper: 🔗Project:show more

Bilawal Sidhu
44,822 次观看 • 2 年前
The architecture of this new world model is one... of the most interesting things I've seen lately: Let me first explain how most world models work: They predict and render one frame at a time. If you are navigating in one of these worlds, and you look left, the model draws whatever looks right in the moment. Every time you change your viewpoint, the model has to imagine what should be there again, so it's very common for these models to "forget" what's in the world. For example, if you put a toy on the table, look away, then look back, the toy might not be there anymore. Tripo AI is releasing its Project Eden model, which works very differently: The model builds the world first, and then renders it based on that map. That map holds the real state of the world: the geometry, every object, where things are, what's already happened. The picture you see on screen gets generated from the map. This architecture flips the whole thing. Now, you get the following: 1. The world stops forgetting. Leave, come back, and the toy is still on the table because it lives in the map, not in the last frame you saw. 2. You can edit the world, and those changes persist for anyone who enters later. 3. Multiple people and AI agents can coexist in the world and see it from different perspectives. This is early research, but it's looking really promising. They just raised nearly $200M across two rounds to build it out. Tripo will be at SIGGRAPH 2026 (July 19–23, Los Angeles Convention Center). If you work in 3D, embodied AI, simulation, or anything spatial, go connect with them there.show more

Santiago
30,244 次观看 • 2 个月前
1/ World models are getting popular in robotics 🤖✨... But there’s a big problem: most are slow and break physical consistency over long horizons. 2/ Today we’re releasing Interactive World Simulator: An action-conditioned world model that supports stable long-horizon interaction. 3/ Key result: ✅ 10+ minutes of interactive prediction ✅ 15 FPS ✅ on a single RTX 4090🔥 4/ Why this matters: it unlocks two critical robotics applications: 🚀 Scalable data generation for policy training 🧪 Faithful policy evaluation 5/ You can play with our world model NOW at NO git clone, NO pip install, NO python. Just click and play! NOTE ⚠️ ALL videos here are generated purely by our model in pixel space! They are **NOT** from a real camera More details coming 👇 (1/9) #Robotics #AI #MachineLearning #WorldModels #RobotLearning #ImitationLearningshow more

Yixuan Wang
129,102 次观看 • 6 个月前
A viral paper "Language Model Represents Space and Time"... recently claims that LLMs learn "world models". As much as I like Max Tegmark's works, I disagree with their definition of world model. World model is a core concept in AI agent and decision making. It is our mental simulation of how the world works given interventions (or lack thereof). A world model captures causality and intuitive physics, telling the agent what is likely and what is impossible. It can and should be used for counterfactual reasoning, i.e. "what ifs": what would happen if I knock over a cup of water? Where would I have been if I had not taken that bus? Yann LeCun Yann LeCun says it well in his position paper ( I quote: "Using such world models, animals can learn new skills with very few trials. They can predict the consequences of their actions, they can reason, plan, explore, and imagine new solutions to problems. Importantly, they can also avoid making dangerous mistakes when facing an unknown situation." The first use of the term World Model in deep policy learning is attributed to hardmaru & Jürgen Schmidhuber: In their seminal paper, an agent masters shooting skills in the popular game Doom (demo below) by learning in imagination, using an internal world model as a "physics simulator". To put in a simple Python math formula, world model learns a function F(s[0:t-1], a) -> s[t:], which takes as input the observed past and current action, and outputs plausible future states. Now the definition of World Model in Tegmark's paper seems to be about predicting GPS coordinates and time eras. I see this as just a classification task with no causal learning and simulation going on. You cannot make meaningful interventions against that model, nor can you optimize any decision making in a closed feedback loop. As for the "space & time neurons", I think they are most similar to the "sentiment neuron" that OpenAI published in 2017: Predicting GPS is conceptually no different from predicting sentiment in my opinion. I don't think their experimental results are wrong - just that their conclusion is on shaky grounds. I welcome any debate! Paper link:show more

Jim Fan
594,014 次观看 • 2 年前
The video smooth zoom on the Samsung Galaxy S26... Ultra is still the closest thing to a professional camcorder experience in the smartphone industry today. In fact, it’s even easier to control than the iPhone. On many other phones, video zooming requires constant finger movement and very precise control. The zoom speed can easily become inconsistent, suddenly speeding up or slowing down. Samsung works differently. You simply hold your finger at a certain position, and the phone continues zooming at a constant speed. The entire process feels extremely stable and linear. It genuinely resembles the powered zoom control of a professional video camera. This logic is fundamentally related to Samsung’s AI slow motion technology. They share the same core foundation: real time control over motion trajectories, speed transitions, and frame interpolation. What you’re seeing here was shot in very windy conditions using Samsung’s Pro Video mode, continuously zooming from 5x to 25x. Aside from some slight stutter during optical lens switching points, the continuous zoom transition within digital zoom ranges is arguably the closest thing to a professional camera currently available on a smartphone. So if the future Samsung Galaxy S27 Ultra really removes the 3x telephoto camera, it could actually improve the video zoom experience further. Fewer optical switching points would theoretically reduce transition jumps and stutters, making the entire zoom range feel even more natural and continuous.show more

Ice Universe
27,250 次观看 • 3 个月前
The faggots showed “routes of Ukrainian drones” to the... Leningrad region — through the territory of Belarus and NATO countries (Picture #1). Now let’s look at reality, confirmed exclusively by their own official sources. Here’s how the air raid alerts developed in Russia on these days: Video for March 23 (Video #1) Video for March 25 (Video #2) Video for March 26 (Video #3) We’ve been working for a long time on our own website — an independent analog of the Ukrainian air alert site, but for the Russian swamps. It shows not only the alerts, but also the work of their air defense. We just managed to finish the recording and history function right before the strikes on the Leningrad region. Interesting detail: as soon as an air alert is declared in the Leningrad region itself, literally a few minutes later it is immediately announced in the Bryansk, Smolensk, Tver, and other regions that lie “on the way” to Leningrad. We do not rule out that the drones could have flown over the eastern districts of Belarus. That option is theoretically possible. However, there is zero confirmation: Belarusian chats in those areas were silent, locals wrote nothing, and no objective evidence (video or photos) has appeared. A simple logical question for all their “experts”: If our drones were not flying over Russian territory, but somehow cruising through NATO countries, then why the fuck did you declare air alerts across the entire country? What exactly were your air defense systems “shooting down” over your own regions? Why did the alerts trigger exactly in the regions that, according to your version, the drones never flew over? And the most interesting part — why does this alert pattern lead perfectly straight to the Leningrad region, where real strikes and fires occurred on exactly those dates? - Night of March 23 — massive attack on the Primorsk port (one of the key oil terminals on the Baltic). A fuel tank was damaged, and the fire burned for over a day. Governor Drozdenco claimed “over 50 drones shot down.” - Night of March 25 — strikes on the Ust-Luga port (Novatek terminal) and Vyborg (a ship was damaged, likely the icebreaker “Purga”, and the port itself). Fire in the port. They claimed “33 drones shot down.” - Night of March 26 — attack on the Kirishi Oil Refinery (KINEF) industrial zone. Fires confirmed by NASA FIRMS satellite imagery. They claimed “21+ drones shot down.” You faggots, when you lie about “routes through NATO”, at least try to coordinate your bullshit with your own official data. The picture is hilarious: alerts across the whole country, air defense working at full capacity, “hundreds shot down” — yet the drones supposedly “never flew over Russia.” Draw your own conclusions, friends. Logic has never been Russian propaganda’s strong suit.show more

DroneBomber
94,375 次观看 • 5 个月前
Robots can now reconstruct 3D scenes in real time... from a single RGB camera. [📍 Projects page + paper] No depth sensor. No retraining. 30 FPS. Researchers at the Imperial College London introduced KV-Tracker, a training-free method that makes heavy models like π³ and Depth Anything 3 fast enough for real-time tracking. The idea is simple. These models use global self-attention, which is powerful but computationally expensive. KV-Tracker caches the key and value pairs from selected keyframes and reuses them for new frames. That cache becomes an implicit scene representation. Result: • Up to 30 FPS • 10 to 15x speedup • Accurate 6-DoF tracking on benchmarks like TUM RGB-D and 7-Scenes • Works with monocular RGB only It also supports object-level tracking with masks and allows saving the KV-cache for later reuse. For robotics, this reduces hardware constraints and moves real-time 3D perception closer to practical deployment. Credit to Marwan Taher (Marwan Taher) at Imperial’s Dyson Robotics Lab and many others who contributed to this! 📍 Save projects page + paper for later: Video: ——- if it matters in AI or Robotics you'll read it here first:show more

Ilir Aliu
53,992 次观看 • 4 个月前
This workflow is perfect for creating short fashion-style cinematic... videos. I simplified the original prompts based on willie’s method, and the whole process is now much faster and more stable: 1. Generate a 3×3 keyframe grid (Nano Banana Pro only) Use this simple prompt: “In a 3x3 grid, show this character in different angles, keep the scene the same, random poses. This is far simpler and more efficient than my old prompts. You can generate multiple times and just pick the keyframes you like most. 2. Extract a high-res keyframe (Super stable trick) Take a screenshot of the keyframe you want from the 3×3 grid, send it back to Nano Banana Pro, and simply say: “Give me a high-resolution version.” This method is much more stable than relying on complex upscaling prompts. 3. Generate the video with Kling 2.5 Turbo Upload the first and last frames to Kling 2.5 Turbo and use this prompt: “The camera very slowly and smoothly lowers on a boom.” From my testing, Kling 2.5 Turbo offers the best balance of stability and cost — other models are either less consistent or noticeably more expensive. 4. Final speed adjustment with willie’s tool (Critical step) Use the tool built by willie to fine-tune the playback speed of each clip. This step is essential for getting that premium cinematic feel. I’ll drop the tool link in the comments.show more

underwood
22,223 次观看 • 8 个月前
We are thrilled to announce a BIG PARTNERSHIP ✨... 🔥 CryptoAutos is revolutionizing the #RWA Space and our token $CNCT is now part of it. You are now able to BUY and RENT cars using the $CNCT TOKEN all around the world 🌎 Welcome to CryptoAutos ✨ The fast, secure and seamless payments platform that connects car buyers with Crypto to a massive marketplace of automotive dealer partners across the globe. Crypto Autos is the home of bringing real-world assets on-chain. Allowing all users to buy, rent and invest in vehicles with crypto. RWAs are all over the news right now, and let’s face it, genuine utility and use case is lacking in Web3. By taking on the traditional slow-moving car market and giving users streamlined access, blockchain technology is already bettering the car industry's infrastructure. $AUTOS powers all this within the Crypto Autos ecosystem. Now we are adding this player into our non stop growing Ecosystem giving CNCT a big chance of being part of a Trillion $ Market. Time to grow ! ⚡️show more

Fixera AI
14,242 次观看 • 1 年前
I am posting after a Long Time on Twitter,... but its to announce a big change. I have started and we are Collecting Egocentric Data at Scale from India, Covering 300 + Commercial Locations 1500 + Households This is the network and base we have built in just past 2 months, as the Robotics companies, VLMs and World Models increase their requirements on Real World Data collection, Human Loops will be keep scaling out capacity. We are on track to collect 1M Hours of Egocentric in the coming 6 months for our clients exclusively. We are maintaining 95% Quality standards across industrial data and running a End - End Operational Management complying all Indian Laws and compensating our partners/operators. (From Environment sourcing, to hardware, Legal contracts, deployment, training, collection, processing) Check our Samples: - Commercial Videos -- Household Videos --- Multimodel ---- Egocentric + Live Audio Narration Links : On the Journey to become #1 India Physical AI Data Partner. We are also building our capacity as annotation and labelling partner for data companies to become end-end partner for companies. A big change from the world of crypto and web3 but physical AI and data space is where i want to build my next venture #Egocentric #EgocentricIndia #PhysicalAI #Robotics #India #AIdata #data #Multimodeldata #worldmodels #VLMs #Humanloops #Egocentridata #Multimodeldatashow more

Shloak
27,867 次观看 • 2 个月前
From International Space Station, we have a unique vantage... point to observe lightning storms from directly above them. Much like from the ground, they appear as rapid, intense flashes that are over as quickly as they started. But, with the help of handheld cameras and a bit of practice (and, admittedly, some luck), we can capture these intense flashes with a very high frame rate then slow them down to see how the lightning propagates. It's lightning, in slow motion. This is one that occurred over Indonesia on 6/28. I shot it for a half second at 1/120 frame rate, so, 60 separate frames. I then compiled those frames into two videos: the first, in real-time, and the second, slowed down so that each frame is a half second to highlight the progression of the flash. Note the incredible burst of light in one frame, which includes what is likely a TLE on the upper left. A TLE, or Transient Luminous Event, is an electrical discharge that happens in the upper atmosphere above lightning storms (if you haven’t seen Ferry Farhan TLE photo from last week, head over to her page). As their name implies, these are transient, or short lived, phenomena – it appeared in only one frame of the sequence, lasting for no more than 1/120th of a second. The last photo is that single frame. Absolutely incredible.show more

COL Anne McClain
82,113 次观看 • 1 年前