For years Yann LeCun has argued that generative video... models can't truly learn physics. DeepMind's Physics-IQ benchmark proved him right, reporting "a striking lack of physical understanding in current generative video models". The best one scored just 29.5%. This CVPR 2026 paper finds the fix in LeCun's own playbook. "Inference-time Physics Alignment" doesn't retrain the generator. It steers a video model's denoising at inference using a reward from VJEPA-2, LeCun's Joint Embedding Predictive Architecture. - Won first place in the ICCV 2025 PhysicsIQ Challenge with a 62.64% score, beating the previous state of the art by 7.42%. - Repurpose VJEPA-2's "surprise" score as a reward, then search and rank multiple candidate denoising trajectories at test time. - Why it matters: It's a neat vindication of LeCun's thesis. The pixel-prediction generator alone doesn't get physics, but the JEPA world model approach he champions supplies the physics the video model lacks.show more

Vai Viswanathan
44,407 views • 3 months ago
This week is already so hot. 🔥 Massive release... from Decart : Lucy 2.0 a World Editing Model running at 1080p, 30FPS in realtime. This is truly exciting, the era of real-time generative reality is here. We are moving from watching AI video to living inside AI video. A breakthrough model capable of transforming the visual world in real-time. Moving beyond offline rendering, Lucy 2.0 delivers high-fidelity 1080p video generation with near-zero latency. Lucy 2.0 literally "redraws" the entire world pixel-by-pixel, while you are watching it. e.g. If you want to be an anime character, it doesn't just put a mask on you. It turns your skin into anime skin, your hair into anime hair, and the lighting in your room into anime lighting. Lucy 2.0 is also trained to stop the generated video from slowly falling apart over time, so the same stream can run much longer without faces and details drifting. So why is this a "Massive Deal"? Traditional AI video-generation model takes a prompt, you wait 10–20 minutes, and the computer "bakes" a video for you. You couldn't touch it or change it while it was happening. But Lucy 2.0 works like a mirror. It happens in real-time (30 frames per second). There is no waiting. You move your hand, the AI character moves its hand instantly. The craziest part isn't the visuals; it's the physics. Usually, AI hallucinations are glitchy—hands merge into faces, walls melt. Lucy 2.0 understands how the world works without being told. It knows that if you take off a helmet, there is hair underneath. It knows that if you splash water, droplets fly. It learned "physics" just by watching millions of videos. The physical behavior you see emerges from learned visual dynamics, not from engineered geometry or explicit physics engines. Their official technical report explicitly states that the model does not use traditional 3D engines, depth maps, or wireframes. It is a "pure diffusion model."show more

Rohan Paul
12,761 views • 8 months ago
Introducing VL-JEPA: Vision-Language Joint Embedding Predictive Architecture for streaming,... live action recognition, retrieval, VQA, and classification tasks with better performance and higher efficiency than large VLMs. • VL-JEPA is the first non-generative model that can perform general-domain vision-language tasks in real-time, built on a joint embedding predictive architecture. • We demonstrate in controlled experiments that VL-JEPA, trained with latent space embedding prediction, outperforms VLMs that rely on data space token prediction. • We show that VL-JEPA delivers significant efficiency gains over VLMs for online video streaming applications, thanks to its non-autoregressive design and native support for selective decoding. • We highlight that our VL-JEPA model, with an unified model architecture, can effectively handle a wide range of classification, retrieval, and VQA tasks at the same time. by Delong Chen (陈德龙) Mustafa Shukor Théo Moutakanni Willy Jade Lei Yu Tejaswi Kasarla Allen Bolourchi Yann LeCun Pascale Fungshow more

Pascale Fung
90,144 views • 9 months ago
Only one Adobe Firefly video prompt from our research... hit production-ready on the first generation. Prompt in comments 👀 It looks you have to direct the camera as much as the model to get the most out of Firefly. Every other video session in our research needed a revision. Revising the first frame is where Firefly broke: >motion failures (temporal flicker, motion continuity, physics in motion) carried 60% of video's issue mix once designers re-prompted. TLDR: in Adobe Firefly, the camera does a lot of the work. Bake it into the first prompt and the video lands. Its by far the best time to be a creative.show more

ben @ CVPR
12,847 views • 4 months ago
A viral paper "Language Model Represents Space and Time"... recently claims that LLMs learn "world models". As much as I like Max Tegmark's works, I disagree with their definition of world model. World model is a core concept in AI agent and decision making. It is our mental simulation of how the world works given interventions (or lack thereof). A world model captures causality and intuitive physics, telling the agent what is likely and what is impossible. It can and should be used for counterfactual reasoning, i.e. "what ifs": what would happen if I knock over a cup of water? Where would I have been if I had not taken that bus? Yann LeCun Yann LeCun says it well in his position paper ( I quote: "Using such world models, animals can learn new skills with very few trials. They can predict the consequences of their actions, they can reason, plan, explore, and imagine new solutions to problems. Importantly, they can also avoid making dangerous mistakes when facing an unknown situation." The first use of the term World Model in deep policy learning is attributed to hardmaru & Jürgen Schmidhuber: In their seminal paper, an agent masters shooting skills in the popular game Doom (demo below) by learning in imagination, using an internal world model as a "physics simulator". To put in a simple Python math formula, world model learns a function F(s[0:t-1], a) -> s[t:], which takes as input the observed past and current action, and outputs plausible future states. Now the definition of World Model in Tegmark's paper seems to be about predicting GPS coordinates and time eras. I see this as just a classification task with no causal learning and simulation going on. You cannot make meaningful interventions against that model, nor can you optimize any decision making in a closed feedback loop. As for the "space & time neurons", I think they are most similar to the "sentiment neuron" that OpenAI published in 2017: Predicting GPS is conceptually no different from predicting sentiment in my opinion. I don't think their experimental results are wrong - just that their conclusion is on shaky grounds. I welcome any debate! Paper link:show more

Jim Fan
594,014 views • 3 years ago
Here's something the physics textbook doesn't dwell on: the... most successful scientific theory ever built runs on 19 numbers that nobody can explain. They're measured. They're plugged in. They work — the Standard Model predicts particle behaviour to extraordinary precision. But ask *why* the electron has the mass it has, or why the strong force is as strong as it is? The theory goes quiet. ISF's geometric framework takes a different approach. It derives its results from the structure of space itself — no numbers fed in by hand. The same framework that produced 0.8412 fm for the proton charge radius (confirmed at 0.8414 fm by independent measurement) needed zero inputs to get there. Send this to someone who loves physics. The 19-numbers problem is one of the field's best-kept open secrets.show more

Nassim Haramein
12,267 views • 5 months ago
Generative 3d environments just became a thing with the... announcement of OpenAI's new video model, Sora. Michael Rublof from took one of those videos, and turned it into a NeRF using Colmap and Nerfstudio. While people are laughing at the topology of generated models, the world is changing around us, and that's exciting and a little scary, but we'll find a way to turn Gen Ai into creative superpowers. I believe in human creativity, in our ability to surprise, to move, to change and to challenge. Here's to the future! #ai #artshow more

Martin Nebelong
288,988 views • 2 years ago
In a masterclass at Sequoia Capital AI Ascent, Jim... Fan laid out the "Great Parallel": how robotics is speedrunning the LLM playbook. 🔹 VLA → WAM: Moving from language-heavy models to "World Action Models" that dream in physics. 🔹 Teleop → EgoScale: Replacing manual data with human egocentric video. 🔹 Simulation 2.0: Using neural simulators like DreamDojo to turn compute into environments. "Our generation was born too late to explore the earth and too early to explore the stars. But we are born just in time to solve robotics." He believes that robots will pass the Physical Turing Test in the coming 2–3 years.show more

Humanoids daily
12,123 views • 4 months ago
As always everyone is blind staring at the progress... of LLMs for coding and chat But meanwhile the new SOTA video model Seedance 2.5 has been slowly rolling out and it's really quite exceptional It's made by ByteDance (TikTok) who of course have lots of training data With just a few reference pics, it can get quite close to how you look IRL and you can do quite professional video shots with just a prompt I'd say it's the first video model that's now at the level of image models with the level of character likeness, cracking that in image models also took about 3 years (2022-2025) Generating 15 seconds takes about 4 minutes I put it live now on Photo AI, you can use it under [ Make video ] from the sidebar with just a prompt and your model selected! So you don't need to take an AI photo first and then turn that into a video! Saves lots of time :D It's more expensive than but I kept the credits the same (30 for 1 video) It also works inside the new video editor and you can make changes in your video with [ Magic edit ] in both the main app and the video editor Also a message for my server guy Daniel Lockyer (it can do voice too and you can even submit a voice sample of yourself, but I didn't here)show more

@levelsio
711,714 views • 1 month ago
we sped up distributed inference by up to 5x... with decentralized speculative decoding. many don't realize that AI models normally generate text one single word at a time, waiting for the network after every word. speculative decoding changes this by using a "guess & confirm" system, similar to autocomplete. how it's done: 1. draft locally (the guess) instead of waiting for the network, a tiny, fast model on your device guesses the next few words instantly, without waiting for the network. 2. confirm remotely (the check) the massive remote model doesn't generate from scratch; it just checks the draft. it looks at the guesses in a batch and says "yes, yes, no." you get multiple words in the time it usually takes to get one. 3. adaptive logic dsd is smart. if the topic is creative, it lets the draft flow loose. if the topic is math or code, it checks more strictly. it balances speed and precision automatically so your inference almost feel instant. find out more: paper: blog:show more

Parallax
45,584 views • 9 months ago
claude fable 5 is live. spawn 5.0 was built... with it: 1,687 prompts, 102 sessions, my job shifted from architecture to judging taste. what we built, each of which would've been at least a month with a whole team on opus: — a from-scratch physics engine (mantle) that rivals rapier, testable in a day as a side thread — clustered froxel lighting: 8 lights to 1,000+ fully dynamic ("it's stupid fast" — creator of threejs) — realtime diffuse GI on webgpu, on your phone (landing in the next update) — million-particle gpu vfx with a shader architecture beyond what unreal and unity ship — mmo-scale netcode plus, for good measure: an (alpha) native ios app. all of it in about a week. all of it running on mobile. and the list doesn't include half of what we shipped — full changelog in the thread below. and the part the benchmarks won't show you: this model has a wonderful personality. it's genuinely funny. i laughed so hard i cried multiple times, mid-physics-rewrite. the genius and the character aren't separate features.show more

jacob
55,515 views • 3 months ago
AI Is Moving Beyond “Generating Videos” — Toward “Generating... Worlds” Over the past two years, AI video models have advanced at an astonishing pace. From Runway and Pika to Sora and Veo, AI-generated videos have become increasingly realistic and more consistent with the physical laws of the real world. Many people believe the next objective is simply to generate videos that are longer, sharper, and more lifelike. But if we take a step back, we can see that the real transformation is not happening in video itself. It is happening in world models. What Is a World Model? In 1943, psychologist Kenneth Craik proposed an idea that would influence artificial intelligence research for decades. He argued that the human brain does not merely react to the outside world. Instead, it maintains an internal model of how the world works. Because we have this internal model, we can predict the outcome of an action before we actually take it. Before crossing a road, we estimate whether a car will pass by. Before catching a ball, we predict its trajectory. These abilities come from continuously simulating the world in our minds, rather than relying entirely on trial and error. This idea later became known by a more formal term: World Model. A world model does not describe a single image or a fixed video clip. It is an internal representation capable of continuously simulating the rules and dynamics of the real world. Why Is AI Research Turning Toward World Models? Because predicting “what comes next” is becoming increasingly central to how AI systems work. Language models predict the next token. Image models predict the next step in the denoising process. Video models predict the next frame. A world model, however, attempts to predict something broader: What should the world look like in the next moment? In 2018, David Ha and Jürgen Schmidhuber proposed in their paper World Models that an intelligent agent could first learn a model of the world, and then use that internal model to plan its actions. The Dreamer series later demonstrated that many complex tasks could be learned by training agents inside an “imagined world.” At the same time, the development of video models such as Sora and Veo led researchers to another realization: A model capable of continuously generating video has already learned, at least implicitly, many of the rules governing the real world. As a result, these two research directions have gradually begun to converge. But Video Is Not Yet a World This is where the distinction is often misunderstood. For a world model to support meaningful real-time interaction, it must solve several critical problems. Most video models today are essentially answering one question: What should the next frame look like? A true world model needs to answer much more: What happens if I take one step forward? If I walk behind a building and then return, will the building still be there? If I suddenly change the camera angle, will the entire space remain consistent? If I enter a command such as: “Summon a dragon.” Will the world respond immediately? In other words, a world model must do more than generate content. It must understand space. It must understand time. It must understand causality. And it must understand interaction. Moving from watching to participating is where the real difficulty of world models begins. World Models Are Entering the Interactive Era One of the latest attempts in this direction is Alaya World, recently open-sourced by Alaya World, or Alaya Lab. Instead of generating a fixed video clip, it generates a world that users can explore in real time. Users can begin with text, an image, or a video, enter the generated scene, move freely through it, and introduce new prompts at any moment during generation. The world responds immediately. According to the publicly released information, Alaya World provides: Real-time streaming generation at 720p and 24 FPS Stable continuous exploration for more than one minute The ability to switch prompts and trigger skills or events during generation Model weights and inference code released under the Apache 2.0 License Training code and datasets planned for future release What makes these capabilities important is not simply the technical specifications. It is that the generated “world” can now support continuous interaction. The official demo shows that users can genuinely control, transform, and explore the generated environment. AI Is Evolving From a Tool Into an Environment Over the past few years, most discussions around AI have focused on content generation. Generating text. Generating images. Generating videos. But world models raise a fundamentally different question: Can AI generate an environment that people can inhabit, explore, and continuously evolve? If the answer is yes, the impact will extend far beyond video generation. Game development, robotics training, embodied intelligence, digital twins, virtual production, and many other fields could be transformed by the development of world models. World models are still at a very early stage. Yet from Craik’s proposal of an internal mental model more than eighty years ago to the emergence of today’s interactive world-generation systems, a clear evolutionary path is beginning to take shape. Perhaps what AI is ultimately learning has never been limited to images, videos, or language. Perhaps it is learning the world itself. References GitHub: Technical Report:show more

雪踏乌云
113,347 views • 2 months ago
THE CASUAL REALISM OF AI IS CROSSING A DANGEROUS... LINE If this popped up on your feed, you would glance at it for half a second and assume it’s just another candid mirror clip saved to someone’s story. A plain room, indoor lighting, a phone case, and a quick pose. There are no glowing sci-fi artifacts or exaggerated physics giving it away. That’s what makes this wave of video models so convincing. For years, mirrors and tight reflective materials like leather or latex were the ultimate stress test for computer graphics. Glass would bend unpredictably, light bounces looked unnatural, and the physical weight of the body never matched the posture. Now, the models are replicating the subtle, imperfect quirks of a phone camera: room glare, natural balance, and casual micro-movements. A year ago, synthetic video was easy to bust within two seconds. Today, generative clips are learning to look just as spontaneous, flawed, and normal as the real world.show more

aitifa uncut
10,723 views • 26 days ago
This broke my mental model of game dev 💀... 2.5 hours → fully playable ‘Worms’ clone. Built with Hermes agent by Nous Research Here’s what made that speed possible: Hermes used ‘Persistent Shell’ mode, which ensured it didn't forget its current folder or active tools. This allowed it to work smoothly, without the distraction of constantly having to recall where it left off last time. To optimize the workflow, the agent moved beyond linear execution and parallelized the workload. It spawned isolated subagents while executing multiple independent tool calls via ThreadPoolExecutor. Like, one subagent wrote Python RPC scripts for the projectile physics while another utilized vision tools for character sprites. When the complex terrain logic required debugging, the agent used filesystem checkpoints and the /rollback command to instantly return to a stable state. To fix UI bugs, it attached to a live Chrome instance via CDP (/browser connect), fixing rendering issues in real-time. The agent’s built-in learning loop was active from the very beginning. By the time the game was finished, this continuous process allowed the agent to autonomously convert the physics logic into a custom skill. This logic is now a permanent plugin file in the agent's plugin architecture, making the physics engine a native capability that the agent can reuse for future projects. Follow War_v3_FINALE.exe for updates!show more

Javier
38,414 views • 6 months ago
Google dropped a new AI paper called LUMIERE. It's... remarkably flexible, supporting video inpainting, image-to-video, AND stylized video generation tasks. Say hello to “space-time diffusion” for video generation! Now what the heck does that mean exactly?! 🌐⏳ → TL;DR it utilizes a “Space-Time UNet” architecture that generates the full duration of the video in one pass, rather than generating distant keyframes and interpolating between them like prior works. Because the computation is done in this “compressed space-time representation” to generate the full clip at once, it's far more temporally consistent. → Another benefit of generating the full video at once is that you can “direct” the video generation, making it easier to hand off to other models/tasks without having to stitch together partial solutions. You can condition generations on additional inputs, meaning you get the full stack of AI video capabilities – from video inpainting to image-to-video and beyond. → New SOTA for AI video generation? User study results in the paper suggest human evaluators preferred Lumiere over Runway Gen-2, Pika Labs, and Stable Video Diffusion in terms of quality, text alignment AND motion. But as always, we need to get hands-on with this tech when Google *actually* decides to ship it. → Could this end up inside YouTube? Y’all know i’m obsessed with blending reality and imagination – so it’s the video inpainting tech I'm most excited about. I really hope this model finds its way into YouTube's Generative AI efforts, and based on their prior announcements and the list of acknowledgments in the paper I think it might! 🤞🏽 Links: 🔗Paper: 🔗Project:show more

Bilawal Sidhu
44,822 views • 2 years ago
Black Forest Labs just announced FLUX 3: a unified... multimodal model for image, video, audio and action prediction. Founded in Freiburg, Germany, one of the few globally significant European companies in the AI sector. The current release is gated early access, not open source. But if FLUX 3 Dev actually ships with usable open weights, this could become one of the most important open-model releases in generative video yet. Especially after the Chinese Minimax H3 release. BFL claims, based on preliminary internal comparisons, that FLUX 3 outperformed models such as Runway Gen-4.5, Luma Ray 3.2, Kling 3 Pro, and Seedance 2.0. However, these figures stem from the early-access phase and do not constitute independent validation. Be that as it may, it's great to see the quality that today's video models can produce at affordable prices. Something that would have been unthinkable a year ago. Really cool release!show more

Chubby♨️
39,869 views • 2 months ago
A loop of AI agents built me a gun.... It just can't make it shoot. Day 11 of building GTA 6 with a loop of agents. Yesterday I said today they'd learn to steal cars and shoot. Today's progress: - 2 weapons in - Pickup system works (we can grab weapons now) The downside: you pull the trigger, nothing happens. The loop is incredible at the overall, but shooting is the main part of the game, so I think we wire this one up ourselves. So here is the progress I'm taking shooting and sitting down to learn the physics myself with agent as support. At the same time J A Z I I takes driving by hand. The agents keep doing what they're good at. They're building out the police system and the economy right now, while I type this. Also spent the weekend in SF at a Tripo3D Donut event learning 3D art.show more

Ziwen
67,385 views • 3 months ago
a physics dropout i know makes $340k at a... quant fund no finance degree. never worked at a bank. got rejected from Goldman twice his edge: he never learned to read charts when he joined, his manager handed him a probability textbook and said "forget everything you think you know about markets" "price is just output," he told me. "we trade the state the market is in" took me weeks to actually understand what that meant every market cycles through a handful of repeating states - compression, trending, volatile expansion, reversion each one has a historically measurable probability of flipping into another. that's it. that's the whole edge you don't predict direction. you identify the current state and bet on the transition with the highest historical frequency built his first model in a weekend. 5 years of futures data, 6 distinct states one state transitioned to a directional move bigger than 1.5 ATR in 72% of cases. reward/risk on those: 4.3 math to do this is in any intro stats course. the framework is literally from 1906. data is free what changed isn't the math - it's that cheap computing finally made it fast enough to run live quant revolution isn't a secret. it's a hundred-year-old idea that finally got hardware retail got moving averages. the physics kids got the actual question Bookmark this before you open another chart they aren't smarter. they were just handed a different textbook from day oneshow more

Livsun
11,040 views • 1 month ago
Qwen3.6 35B A3B can't fill out a paper form... on its own. But give it NVIDIA's LocateAnything-3B — the #1 trending model on HuggingFace — as its eyes, and the two small models get it done together. (The test: place each element at the right pixel position on a blank form image, not type into a field.) Setup: > Qwen is the brain (main model), LocateAnything is the eyes (helper model acting as a tool). > I gave Qwen a new tool: ask "where's the email field?" and LocateAnything returns the exact x, y, width, height. > The blue boxes on the screen are its detections. Look how tight they are — it nails every field. Result: > Qwen3.6 35B A3B + LocateAnything-3B: form completed, all info correct. > Name, DOB, ID, gender, marital status, nationality, email, phone, address, postal code: all landed in the right field areas. > Character-box alignment still a touch loose, but every value is where it belongs. > 9m10s, 224.5k input, 24.3k output, 21 turns. Why it matters: > Qwen alone can't finish this test. Bolt on a 3B model that does exactly one thing > locate > and suddenly it can. > A combination of small models can do the work of a single large one.show more

stevibe
150,847 views • 4 months ago
If you typically stream with a 3D model, I... highly recommend that you pose your model when you’re not doing full body mocap instead of letting it stay in the stiff generic pose! It makes a huge difference in how your energy is conveyed (ᗒ⩊ᗕ)⸝ި ʕᦏ⌎ I’m only saying this because I noticed this a lot with my own 3D kids, but here’s a comparison showcase: ← default pose & default pendulum physics in Warudo → custom poses & slightly adjusted pendulum physics If you find it hard to pose within Warudo by using bone offsets or none of the existing poses in Warudo vibe with you, you can make your own like me! I made my custom poses in “VRM Posing Desktop” (this is the BEST vrm posing app I’ve ever used in the past 3 years) and exported them as Unity anims then dropped them into Warudo’s animation folder! You can then make a simple blueprint in Warudo to toggle between poses and make your model look more alive! It should fit especially well for just chatting streams 🙂↕️✨show more

𝗞𝗔𝗥𝗜𝗛𝗔 🌘🍀 𝟯𝗗 𝗔𝗿𝘁𝗶𝘀𝘁 ☻
16,214 views • 7 months ago