Загрузка видео...

Не удалось загрузить видео

На главную

What if we can take a few photos and turn it into an interactive physically realistic virtual world🌎? 📢Introducing AdaVoMP (ICML 26): generates volumetric physics fields at high spatial resolution making objects interactive and deformable🦾 🦣16^3x higher res (1024^3) ⚡️more accurate

15,334 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 6

Фото профиля rishit dagli
rishit dagli3 месяцев назад

3D assets usually have geometry and appearance like a Gaussian Splat you can capture📸, but not the physics needed for simulation To make them move realistically, we need physics fields throughout the volume. This requires generating the internals of the 3D assets too.

Фото профиля rishit dagli
rishit dagli3 месяцев назад

This continues on our VoMP (ICLR 26, line of work on fast accurate physics from any 3D object so they can be made interactive but: ⚡️higher resolution 🤖new learned adaptive structure 🗺️increase test time compute to increase generation resolution

Фото профиля rishit dagli
rishit dagli3 месяцев назад

We first introduce a new representation, SAV to encode input and generate over 🧱Large homogeneous regions can stay coarse. 🔬Only complex material boundaries and heterogeneous regions are recursively refined. 🎯So the model spends resolution where the physics actually changes.

Фото профиля rishit dagli
rishit dagli3 месяцев назад

2 new models: 1. Adaptive Geometry Transformer: encodes high-res 3D inputs from multi-view DINOv3 features🧊 2. Adaptive Material Generator: autoregressively generates a material tree from 1³ → 2³ → 4³ ... up to 1024³ by running the generator in a loop🌀 and it scales well📈

Фото профиля rishit dagli
rishit dagli3 месяцев назад

🤖Plug into robot evaluation frameworks! 🚀On hard test set (1024^3) we are over 25% and 30% more accurate on YM and density error than VoMP ⚡️at low-resolution (64^3) our structure uses only 9.14% of occupied voxels that VoMP operates over making our structure more efficient

Фото профиля rishit dagli
rishit dagli3 месяцев назад

📜 Paper: 🌐 Project Page: None of this would have been possible without the invaluable support from the team @DonglaiXiang @vm2358 @xuningy @gavrielstate @diwlevin @_shumash at @nvidia @NVIDIAAI @UofT @UofTCompSci And thanks to Gilles Daviet, Jean-Francois Lafleche, @BPerschall, Katherine Cheung, Ruchik Thaker, Andre Pradhana, @AnkaHeChen, Anita Hu, Charles Loop, @Caenorst, @frncswllms, @zhaohexu2001, Ken Museth

Похожие видео

50 NEW internet rabbit holes you can spend hours exploring👇 No repeats. Just websites worth getting lost in. 🌐 1. — Spin the globe and listen to radio stations around the world 2. — Watch Earth's winds, oceans, and atmosphere move 3. — Explore weather patterns across the planet 4. — See fishing activity across the world's oceans 5. — Explore volcanoes and their history 6. — Explore satellite imagery of Earth 7. — Explore a 3D virtual globe 8. — Explore climate science and visualizations 9. — Dive into the science of our oceans 10. — Explore the weirdness of the internet 11. — Test your geography skills with location guessing games 12. — Travel the world without leaving your chair 13. — Discover locations using three-word addresses 14. — Explore street-level imagery from around the world 15. — Browse crowdsourced street imagery 16. — Explore interactive maps and geographic stories 17. — Explore historical geography through maps 18. — Discover old maps from around the world 19. — Explore geographic coordinates and locations 20. — Discover places and details on a world map 21. — Explore planets, asteroids, and spacecraft in 3D 22. — Discover worlds beyond our solar system 23. — Explore solar activity and space weather 24. — Track satellites and celestial objects 25. — Explore the night sky and astronomy events 26. — Explore a virtual universe 27. — Dive into the planets and moons of our solar system 28. — Experiment with interactive science simulations 29. — Explore interactive science and math explainers 30. — Explore physics concepts through linked explanations 31. — Ask questions and explore computational knowledge 32. — Learn math and science through interactive problems 33. — Explore interactive explanations of complex ideas 34. — Discover interactive science and math experiments 35. — Explore the strangest corners of science 36. — Discover scientific papers and research 37. — Search open-access academic research 38. — Dive into millions of academic articles 39. — Explore scientific preprints 40. — Fall down the mathematics rabbit hole 41. — Explore fascinating historical objects and stories 42. — Discover rare books and historical collections 43. — Explore photos, maps, and archives 44. — Browse the Louvre's art collection 45. — Explore artifacts from human history 46. — Discover objects from museums across Scandinavia 47. — Explore the world's music database 48. — Discover movies and endless film rabbit holes 49. — Explore sounds for focus, relaxation, and sleep 50. — Click a button and discover a random website Save this. Your next internet rabbit hole is waiting. 🔖 Follow Kaizen 黒 for more useful websites, AI tools & tech resources.

Kaizen 黒

31,888 просмотров • 11 дней назад

Most AI world models can generate beautiful scenes. Keeping those scenes alive for an hour without falling apart is the real challenge. That's what caught my attention about LingBot-World 2.0 (LingBot-World-Infinity) from Robbyant Instead of chasing longer videos, it focuses on something much harder: persistent, interactive worlds that stay coherent while you explore. A few highlights: • Generates worlds from a single frame and continuously responds to live user actions through a causal world model. • Streams stable 720p at 60 FPS in real time. The team reports a continuous 60 minute stress test across 20 different scenarios with no noticeable visual degradation. • Uses a Brain-Cerebellum co-simulation framework where a VLM plans events while the video model turns them into consistent world evolution. • Pilot and Director Agents help drive character behavior and introduce new objects and events. • Open sourced with a 14B flagship model, while the paper also describes a lightweight 1.3B version for a single consumer GPU. There is also an online interactive demo. The biggest takeaway? We're moving beyond AI that generates clips. We're getting closer to AI that generates living, evolving worlds you can actually interact with. And that feels like a much bigger shift than another jump in video quality. Explore more: 💻 Github: 🤗 Weights: 🌐 Website-with videos you can use : 🎮 Try it online: #Robbyant #LingBot #WorldModel #EmbodiedAI #OpenSource #Robotics #ad

Alif Khan

84,547 просмотров • 2 месяцев назад

One of the things I’m most excited about in our recently announced partnership with Niantic Spatial 🌎, is how clearly it shows what becomes possible when world-class reconstruction technology is paired with a new kind of imagery infrastructure. At a high level: Spexi drone pilots capture imagery, and Niantic Spatial turns it into incredible city-scale reconstructions. But the real unlock is the infrastructure behind that capture. At Spexi, we’ve built what we believe is the world’s first fully standardized drone imagery infrastructure called LayerDrone. Anyone with a compatible drone and the right credentials can contribute. No building flight plans. No estimating overlap. No adjusting camera settings in the field. Pilots simply get within visual line of sight of a Spexigon, open the Spexi app, press “Fly,” and the drone autonomously captures the 25-acre area to our standard. That standardization means imagery can be collected consistently, affordably, and repeatedly across cities, one Spexigon at a time (we have now captured over 225,000 of them). That is what makes living digital twins possible, dynamic representations of the physical world that can be updated as the world changes. Niantic Spatial’s city-scale Gaussian splats show what becomes possible when the right pixels go into the system. As physical AI advances, those pixels matter even more. Robots, drones, vehicles, maps, and spatial intelligence systems will all need current, high-resolution data about the real world. And as you can see below.. the results are not just beautiful, but real, measurable reconstructions of the physical world, one Spexigon at a time!

Alec Wilson

10,695 просмотров • 3 месяцев назад

Introducing /visual-plan - a skill to generate rich, visual plans for Claude Code and Codex. Plan mode in Claude Code is incredible. But I always find my eyes glazing over when it gives me this huge markdown essay in my terminal. I found I can make much better visual plans with reusable components. So I made a skill called `/visual-plan`. It generates plans as MDX with visual, interactive components. Diagrams, interactive API specs, schema design changes, annotated code, and even pan and zoomable wireframes. So for any UI work, you can look at a wireframe first, comment on it, iterate, and then have the agent work. I’ve found this to be a much more intuitive interface for reasoning about what the agent is doing. It’s somewhat inspired by that popular post about how HTML is better than Markdown. But HTML can be slow and verbose to write. And it doesn’t look good checked into a repo. This has really made me feel like humans and engineering are entering a new abstraction phase, where we reason about things at the plan level. As long as the plan is good, agents are getting more and more reliable at executing on it. Almost to the degree that we trust the C compiler to compile to assembly reliably. Plans are the new intermediate representation. I also made a skill for the reverse of this, called `/visual-recap`. After the agent works, it gives you a recap of everything it did. Same idea: wireframes, interactive API specs and diffs, schemas, annotated code, etc. So now when you’re reviewing what the agent did for you, or looking at a pull request of somebody else’s code, you can see a visual recap instead of just reading a wall of text. It’s all free and open source. You can find it on my GitHub. Will link to it in the reply because we all know how dumb these algorithms are with links.

Steve (Builder.io)

126,165 просмотров • 3 месяцев назад

Today at Stanford, Fei-Fei Li (Fei-Fei Li),Cofounder/CEO World Labs, gave one of the clearest explanations I’ve heard of what a World Model really is. She broke it down into three layers: 1️⃣ Rendering — What does the world look like? This is where most of today’s video generation models operate: generating increasingly realistic and beautiful pixels. The question is: Can AI generate what the world looks like? The primary consumer is humans. 2️⃣ Simulation — How does the world actually work? Fei-Fei gave a simple example: “How will this bottle move? If I pour the water out, how will the water flow?” This goes far beyond generating something that looks realistic. The model needs to understand physics, spatial relationships, cause and effect, and how the world changes over time. The consumers are both humans and machines. 3️⃣ Planning — What should happen next? This is where things get really interesting. AI doesn't just render the world or simulate what might happen. It uses its understanding of the world to decide: What should I do next? At this layer, the primary consumer is the machine itself. And this connects directly to two enormous opportunities: Autonomous driving and robotics. The progression is powerful: Rendering → Simulation → Planning The real promise of World Models isn't simply generating better videos. It's building AI that can understand the world, predict what happens next, and ultimately take intelligent action in the physical world.

PaulFang

11,720 просмотров • 1 месяц назад

🚀 Announcing Echo — our new frontier model for 3D world generation. Echo turns a simple text prompt or image into a fully explorable, 3D-consistent world. Instead of disconnected views, the result is a single, coherent spatial representation you can move through freely. This is part of a bigger shift in AI: from generating pixels and tokens to generating spaces. Echo predicts a geometry-grounded 3D scene at metric scale, meaning every novel view, depth map, and interaction comes from the same underlying world — not independent hallucinations. Once generated, the world is interactive in real time. You control the camera, explore from any angle, and render instantly — even on low-end hardware, directly in the browser. High-quality 3D world exploration is no longer gated by expensive equipment. Under the hood, Echo infers a physically grounded 3D representation and converts it into a renderable format. For our web demo, we use 3D Gaussian Splatting (3DGS) for fast, GPU-friendly rendering — but the representation itself is flexible and can be easily adapted. Why this matters: consistent 3D worlds unlock real workflows — digital twins, 3D design, game environments, robotics simulation, and more. From a single photo or a line of text, Echo builds worlds that are reliable, editable, and spatially faithful. Echo also enables scene editing and restyling. Change materials, remove or add objects, explore design variations — all while preserving global 3D consistency. Editing no longer breaks the world. This is only the beginning. Echo is the foundation for future world models with dynamics, physical reasoning, and richer interaction — environments that don’t just look right, but behave right. Explore the generated worlds on our website and sign up for the closed beta. The era of spatial intelligence starts here. 🌍 #Echo #WorldModels #SpatialAI #3DFoundationModels Check it out:

SpAItial AI

177,073 просмотров • 9 месяцев назад

Review: The Witcher 3 Remastered is a massive overhaul to an all time classic RPG - but I have some thoughts. ▫️ Remaster looks sharper, with much higher quality assets across the board ▫️ Ray tracing has been massively improved. Bounce lighting is way better, water reflections finally look amazing, and the shadows are sharper and more accurate ▫️ I do think the lighting needs another pass in a few key areas like Kaer Morhen ▫️ The original has more contrast while the remaster has an overall brighter look (which some may prefer) ▫️ While the remaster is probably more 'accurate' overall, it does lose some of that stylized look from the original ▫️ Performance is definitely much heavier, especially with ray tracing ▫️ The dynamic camera during combat is a big improvement ▫️ Dodging feels smoother and animations are a lot snappier and more responsive ▫️ Meditation is a lot better, it happens in realtime so you can cancel it exactly at the right moment ▫️ Roach handles better and doesn't get stuck on objects quite as much ▫️ Skill tree system has been overhauled ▫️ Transmog has been added (unlocked at Skellige) ▫️ Photo Mode has been greatly improved ▫️ Movement feels better and more responsive There are a lot more changes, but overall it's a nice improvement to an already amazing game. Nothing groundbreaking but some pretty good changes for a free update! The visuals are absolutely stunning in places like Skellige and Toussaint. And it does look much sharper and higher resolution overall. But the lighting needs a few more tweaks in the areas I mentioned to retain the look and feel of the original. It certainly doesn't look bad, just not quite as punchy. Overall, I really enjoyed my time with it and was immediately hooked back into the world of The Witcher. Can't wait for Songs of the Past next year, and of course, The Witcher IV! Final Score: 9.5/10 Thanks to CD Projekt Red for providing code for review. #TheWitcher #TheWitcher3 #Witcher3Remastered

KAMI

468,255 просмотров • 7 часов назад