✨ Any static 3D assets ➡️ 4D dynamic worlds.... Introducing CHORD, a universal framework for generating scene-level 4D dynamic motion from any static 3D inputs. It generalizes surprisingly well across a wide range of objects 🤯 and can even be used to learn robotics manipulation policy 🤖! Project page: Dive deeper in a 🧵: 1/nshow more

Chen Geng
43,493 次观看 • 7 个月前
DroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from... In-the-Wild Drone Imagery Abstract: Drones have become essential tools for reconstructing wild scenes due to their outstanding maneuverability. Recent advances in radiance field methods have achieved remarkable rendering quality, providing a new avenue for 3D reconstruction from drone imagery. However, dynamic distractors in wild environments challenge the static scene assumption in radiance fields, while limited view constraints hinder the accurate capture of underlying scene geometry. To address these challenges, we introduce DroneSplat, a novel framework designed for robust 3D reconstruction from in-the-wild drone imagery. Our method adaptively adjusts masking thresholds by integrating local-global segmentation heuristics with statistical approaches, enabling precise identification and elimination of dynamic distractors in static scenes. We enhance 3D Gaussian Splatting with multi-view stereo predictions and a voxel-guided optimization strategy, supporting high-quality rendering under limited view constraints. For comprehensive evaluation, we provide a drone-captured 3D reconstruction dataset encompassing both dynamic and static scenes. Extensive experiments demonstrate that DroneSplat outperforms both 3DGS and NeRF baselines in handling in-the-wild drone imagery.show more

MrNeRF
21,386 次观看 • 1 年前
DimensionX: Create Any 3D and 4D Scenes from a... Single Image with Controllable Video Diffusion TL;DR: Create 3/4DGS from Video Diffusion Note: Some first inference code released (not all yet). Contributions (cited): • We present DimensionX, a novel framework for generating photorealistic 3D and 4D scenes from only a single image using controllable video diffusion. • We propose ST-Director, which decouples the spatial and temporal priors in video diffusion models by learning (spatial and temporal) dimension-aware modules with our curated datasets. We further enhance the hybriddimension control with a training-free composition approach according to the essence of video diffusion denoising process. • To bridge the gap between video diffusion and real-world scenes, we design a trajectory-aware mechanism for 3D generation and an identity-preserving denoising approach for 4D generation, enabling more realistic and controllable scene synthesis. • Extensive experiments manifest that our DimensionX delivers superior performance in video, 3D, and 4D generation compared with baseline methods.show more

MrNeRF
17,062 次观看 • 1 年前
WeatherEdit: Controllable Weather Editing with 4D Gaussian Field Contributions:... 1. Based on our analysis of weather editing characteristics, we introduce WeatherEdit, a comprehensive and efficient framework for realistic and controllable weather generation. Compared with existing methods that focus on either background editing or static weather effects, a progressive 2D-to-4D transformation process in WeatherEdit enhances adaptability across a wider range of scenarios. 2. We introduce an all-in-one adapter to enable a diffusion model for multi-weather (snowy, rainy, and fog) synthesis, along with a Temporal-View attention to ensure consistent editing across multi-frame and multi-view. 3. We design a 4D Gaussian field for weather particle modeling, enabling plausible simulation of raindrops, snowflakes, and fog with controllable severity. 4. We demonstrate WeatherEdit’s effectiveness in generating realistic, consistent, and controllable weather effects in 3D driving scenes, showcasing its applicability to real-world scenarios.show more

MrNeRF
10,691 次观看 • 1 年前
Imagine making 2D concept art for a game world... –pressing a button – and suddenly you can walk around an interactive 3D world. That's what Google DeepMind's new paper Genie 2 can do – simulate virtual worlds, including the consequences of any action (e.g. unlock door, jump, swim etc). Right now Genie 2 can generate consistent worlds for up to a minute. And this world model seems to generate larger 3D worlds than what World Labs showcased yesterday. Plus they're dynamic vs. static worlds – the foliage moves in the wind, the water ripples etc. Not quite ready for prime time, but promising on two fronts: 1. For game developers: enabling rapid prototyping of interactive experiences straight from concept art 2. For AI research: providing unlimited, diverse 3D environments for training and testing AI agents The race for building the biggest, baddest world model is very much on. Meanwhile, all I can think is "if only Stadia was still around!"show more

Bilawal Sidhu
71,398 次观看 • 1 年前
I am blown away 🤯. Check this out! CameraCtrl... II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models TL;DR: "To enable broader exploration of dynamic scenes, our model can generate new video clips of the same scene based on previously generated content and user-provided camera trajectories. This approach maintains dynamic capabilities, accurate camera control, and scene consistency throughout the extended exploration." "Our model enables precise camera control across diverse scenarios while preserving dynamic scene elements, e.g." "Our method can generate videos with strong 3D consistency, which enables high-quality 3D reconstruction using the camera-controlled videos." Contributions: 1) A systematic data curation pipeline for constructing a dynamic video dataset with camera trajectory annotations; 2) A lightweight camera control injection module and corresponding training strategy that preserves dynamic video generation capabilities while adding camera control effect; 3) A clip-wise autoregressive generation recipe that enables extended range exploration of generated scenes.show more

MrNeRF
12,633 次观看 • 1 年前
Today, we released Lyra 2.0, a framework for generating... persistent, explorable 3D worlds at scale, from NVIDIA Research. Generating large-scale, complex environments is difficult for AI models. Current models often “forget” what spaces look like and lose track of movement over time, causing objects to shift, blur, or appear inconsistent. This prevents them from creating the reliable 3D environments required for downstream simulations. Lyra 2.0 solves these issues by: ✅ Maintaining per-frame 3D geometry to retrieve past frames and establish spatial correspondences ✅ Using self-augmented training to correct its own temporal drifting. Lyra 2.0 turns an image into a 3D world you can walk through, look back, and drop a robot into for real-time rendering, simulation, and immersive applications. ➡️ Learn more: 📄 Read the paper:show more

NVIDIA AI Developer
437,944 次观看 • 4 个月前
NVIDIA finally released Neuralangelo's source code! The model can... turn videos from any device into detailed 3D structures, fully replicating buildings, sculptures, or other real aworld objects or spaces virtually. Here's how it works: A model utilizes a 2D video with multiple angles of an object or scene. I selects frames from different viewpoints to understand depth, size, and shape. The AI creates an initial 3D representation, similar to a sculptor shaping a subject. The render is optimized to enhance details, like a sculptor refining texture. The outcome is a 3D object or scene suitable for virtual reality, digital twins, or robotics.show more

Lior Alexander
478,069 次观看 • 3 年前
What if you could turn a single 360° photo... into a production-ready Isaac Sim environment in minutes? That's exactly what we did here. Using World Labs' Marble and an Insta360 X5 capture (rotating on top), we generated a complete navigable 3D environment and populated it with Lightwheel Sim Ready assets (bottom view). The result? A fully interactive scene in Isaac Sim, ready for sim2real testing,. Navigation, manipulation, or any robotics task you need to validate. What used to take weeks of manual 3D modeling and asset placement now takes minutes. Capture once in the real world, simulate everywhere in your training pipeline. This is the future of robotics development with world models. NVIDIA Robotics NVIDIA Omniverse #Sim2Real #Robotics #Simulationshow more

Jonathan Stephens
46,643 次观看 • 8 个月前
Loving how this turned out! IronSight turns Meta Ray-Ban... clips (from the range) into 4D reconstructions you can replay from any angle -- including an AR view that sees targets straight through walls. It 3D tracks both runs, auto locates every target, and scores hits vs misses using audio cues + Gemini for multimodal reasoning. Full breakdown coming to the channel. The test below is where this started, and then Fable showed up and I blitzed through my whole roadmap in a few days.show more

Bilawal Sidhu
26,530 次观看 • 2 个月前
I’ve been craving a video editor with motion baked... in, designed to work with my own AI agents. Claude, Codex, whatever I want to use. I don’t want to be blocked by someone else’s credit system for every little interaction. I also wanted a Figma-like canvas/editor for when I want to jump in and tweak any little detail. So we built it: Video & audio editing: cut, zoom, speed up, add B-roll. Motion, animation, 3D, custom shaders & effects, custom 3D models, dynamic components & templates, audio and beat matching. And it’s all programmable. Design your own assets directly on the canvas, bring them in from Figma or the web, or just let the agent find them, mock them up, or record what it needs automatically. You can even drop in your screen recordings and just let loose. Supercut users: yes, it’ll have first-class support for your recordings, but it’ll work with anything. I’m going to be dropping a lot more examples and tutorials, so follow along. DM me for early access if you’re willing to give feedback. We’ll open it up to everyone very soon.show more

Neil
23,042 次观看 • 1 个月前
$KNDX 🤖 Theres 3 big narratives that are sending... coins left right and centre rn. 🚀 #AI, #Gamefi, & #NFTs 🔹Theres 50% mindshare for #AI. 🤖 🔹#GameFi mcap is hitting ATH's with #OfftheGrid, $XBG and $SUPER making spectacular moves. 🎮 🔹NFTs and the #Metaverse are making a strong comeback with $APE up 100% over the weekend. 🐵 What if there's a project that touches all these trending narratives with groundbreaking technology to disrupt all 3 of them? 🔥 💡- That's where $KNDX comes in. -💡 Kondux is a cutting-edge Web3 SaaS platform, combining NVIDIA’s Omniverse, AI, Blockchain, and dynamic NFTs to revolutionize secure asset management across industries. 👏 Their flagship product, kNFTs, are 3D digital assets usable across Metaverse and Gaming platforms, AR/VR/XR environments, and manufacturing applications. Kondux’s scalable model opens new revenue streams by enabling effective digital asset monetization. 💰 Kondux is the first Web3 project to integrate VFX pipelines with NVIDIA’s Omniverse and bringing it onto the Blockchain. ⛓️ It is also the only Web3 project with a *Select Status Partnership* with NVIDIA, operating under NVIDIA NDAs and working with them directly for more than 2 years. About their NVIDIA Integrations: 🤖 🔹There are three areas of the Kondux tech stack that coincide with three divisions of NVIDIA: 📡GDN (Graphics Delivery Network, the backbone of GeForce Now) 💡Omniverse for 3D aspects such as, geospatial data, real world physics, lighting, and raytracing 🤖NVIDIA AI Foundation, which covers many aspects of #AI, including inference and deployment scaling. The convergence of all these components lie within .USD file format . 🔹 They are the first blockchain project to integrate NVIDIA’s Omniverse Cloud and Graphics Delivery Network (GDN) to provide high-quality 3D content accessible on any device without requiring high-end hardware. 🔹 This setup streamlines content management, democratises access to resource-intensive 3D content, and enables real-time interaction with 3D NFTs. Now, I haven’t seen any crypto project so deeply connected with NVIDIA and NVIDIA technology. GDN is a HUGE competitive advantage. With it, the need for #GPU’s basically goes out the window. 🤯 Now lets take a look at some of the other main features... 👀 OpenUSD (Universal Scene Description): 📽️ 🔹 Kondux is leveraging USD technology, developed by Pixar and used by Meta, Apple, Microsoft and other industry leaders to enhance 3D graphics and interoperability within its creative ecosystem. 🔹 Originally created for high-end film production, USD now supports a variety of applications, including gaming and virtual reality, making it a key asset for Kondux. kNFT's: 🎨 🔹 Kondux is pioneering a new category of NFTs known as kNFTs, which aim to redefine NFT utility through innovative features. 🔹 A standout feature is the upgradeable aspect provided by Kondux DNA, allowing kNFTs to transform and combine with other NFTs, creating limitless possibilities in art, gaming, and music. 🔹Through the Kondux AI portal it will be possible to communicate with kNFTs. They can learn and adapt. This AI technology is revolutionary because it makes human to kNFT interaction possible, turning it into a unique, personalized experience. Check out the clip of kNFTs in Unreal Engine 5 gameplay below. 👇 Kondux is a very obvious utility play with huge upside because it’s multi narrative. 📈 It's seriously groundbreaking stuff that they’re about to launch. 🚀 After speaking with the team there’s no doubt in my mind this will do crazy big numbers in the next months. 🤑show more

Altcoin Miyagi🇯🇵
17,323 次观看 • 1 年前
✨ I can now generate 3d assets for my... drone sim at directly from Cursor (sponsor of #vibejam) I need buildings that you'd see in a war torn city, like warehouses in ruins, broken down abandoned houses, bombed out bridges etc. Nano Banana Pro or 2 can generate them really well and then you can put them in an image-to-3d model and you get a GLB or FBX That one you can then import into your Three.js game, the models might be big though, in my case like 16MB, so I ask it to compress it and make it more low poly so it loads fast ThreeJS then loads the individual GLBs on page load and puts them in my drone sim somewhere randomly, I think I should remove some of the grass and match the sandy color of the ruins though to make it fit in moreshow more

@levelsio
134,777 次观看 • 5 个月前
🔥 VIDU Multi-Entity Consistency Give Vidu 2/3 images and... it’ll turn them into a video—it’s pure magic! ✨ Your own characters interacting with objects and in the exact environment you want! Ads, movies… endless possibilities, and this is just the beginning! Thanks @Viduforhuman The future is a carrot! 🥕 Plus, how about grabbing any frame from a Vidu-generated video "from scratch" and throwing it into another AI video or image tool to push your project even further? For now, check out the comment below: I scaled up a frame with Magnific.ai and fed it into Runway to create a dynamic shot using full camera control. But fingers crossed I can soon use #ReCapture by Bisho & team to generate new shots from the same video!show more

Hungry Donkey 🥕
37,561 次观看 • 1 年前
You can't 3D reconstruct glass from images... ...WRONG! Thanks... for video diffusion, now just about anything is possible! Introducing...Diffusion Knows Transparency (DKT) Transparent and reflective objects usually break robot vision and photogrammetry pipelines because they don't follow the "solid object" rules standard cameras expect. DKT is a new AI model that repurposes the "internal physics engine" found in video generation models to solve this problem. Researchers took a massive video diffusion model (WAN) and fine-tuned it using a custom-built synthetic dataset to turn it into a high-precision depth sensor. To train the AI, they built the first massive synthetic video library of transparent objects, 1.32 million frames of perfectly labeled glass and metal objects in motion. Without ever seeing a "real" labeled video of glass during training, the model (DKT) outperformed all previous specialized systems on real-world benchmarks (ClearPose, DREDS). They created a "lightweight" 1.3B parameter version that runs fast enough (0.17s per frame) to be used on actual robot hardware. Two reasons I find this project important: 1. It further proves that synthetic data will be essential for training the next generation vision models. 2. In real-world robotic tests, using DKT's depth maps nearly doubled the success rate of robot arms trying to pick up objects on tricky reflective or translucent surfaces. At home robots will need to interact with these types of objects on a daily basis. Check out the project page here: Code is LIVE! #Computervision #Robotics #AIshow more

Jonathan Stephens
17,712 次观看 • 8 个月前
Thrilled to unveil Youmio, our new brand identity that... represents the next evolution of what we’ve been building. Agents are the biggest technological leap since the internet, destined to transform crypto, games, and entertainment. With Youmio, we are shaping the agentic era, where agents learn, play and entertain in revolutionary ways. 🚀 So far, 2D entertainment and social media agents dominate the market. 3D agents are rare, requiring advanced AI and game engine skills. Yet 3D agents, especially those in game engines, unlock groundbreaking opportunities. Time to unleash them. Youmio empowers anyone to create and deploy valuable agents that are on-chain, cross-platform and ready for 3D worlds. Here’s how: ⭐️ Youmio Agents Youmio Agents lets anyone design and personalize 3D agents, equip them with powerful agentic capabilities, interact with them in unique ways and trade seamlessly within a cross-platform browser experience. 🕹️ Youmio Worlds Previously known as Today The Game, Youmio Worlds is a petri dish AI simulation where users build & co-inhabit beautiful, living 3D worlds with autonomous agents. Build dynamic worlds where players interact with intelligent agents, manage resources and participate in a player-agent marketplace. Ancient and Mythic Seeds are the most powerful entry points into the Youmio Worlds ecosystem, generating rare and beautiful worlds that unlock unique opportunities. 📡 Interoperable 3D Agents With Youmio, you’re not limited to our ecosystem. Using our API, developers can integrate Youmio agents into other experiences built in Unity and Unreal. On top of this, agents from other frameworks can also join Youmio, creating a truly interconnected metaverse. 🎭 Welcome to Limbo Meet Limbo, the first AI agent built using Youmio tech. Paired with the power of Youmio Worlds, we’re creating the Limboverse - a unique AI Big Brother setting where Limbo and your favorite and most valuable agents coexist in an ever-evolving, narrative-driven environment that you, the audience, will shape. $LIMBO is the most powerful entry point into the Limboverse and will be stakable on the Youmio Agents platform for unique rewards. Thanks for reading everyone and thanks for being on this amazing journey with us.🌱show more

Youmio
126,460 次观看 • 1 年前
With Hunyuan3D World Model 1.0 now released and open-sourced,... we're excited to showcase the technical highlights behind this impressive innovation: ✅360° Panoramic Generation: Creates complete, immersive “world scenes”, far beyond localized views. ✅Explorable 3D Scene Generation: Generates diverse, spatially consistent 3D worlds from text/image for truly immersive exploration. ✅Interactive/Editable: Achieves separation of foreground objects, background terrain, ground, and sky, for seamless secondary editing. ✅Exportable Mesh: Generated scenes can be exported as 3D meshes for direct import into mainstream game engines and modeling software. ✅Industry-Leading SOTA Evaluation: Surpasses state-of-the-art open-source models in generation quality. As the industry's first open-source model for physical simulation and explorable world generation, Hunyuan3D World Model 1.0 aims to foster a collaborative community ecosystem with developers and enthusiasts. ✨ Try it now: 🤗 Hugging Face:show more

Tencent Hy
23,240 次观看 • 1 年前
How a 22-year-old developer built a full 3D Jet... Ski racing game in just 40 minutes with zero manual coding He used Claude Opus 5 to generate physics, WebGL 3D graphics, HUD, and audio in a single prompt and turned single-prompt gamedev into a high-margin income stream. Costs: $423 He launched a single-prompt generation workflow that built the entire HTML5 project from scratch: Top layer: A Three.js and WebGL rendering pipeline dynamically creates 3D water physics, real-time wave dynamics, dynamic lighting, and jet ski fluid mechanics, all written autonomously inside one output file without external frameworks. Bottom layer: The Claude Opus 5 engine processed a massive 690-million-token context window to generate the complete gameplay logic, collision handling, dynamic sound generation, controls, and UI layout directly from a detailed initial system prompt. The trend of single-prompt 3D game creation is rapidly exploding across media and indie development. The author monetizes this tech stack through three main channels: 1. Viral Content & Media Systems: Short-form breakdown videos driving massive reach, monetized via promo placements, prompt-pack access, and private developer communities. 2. Rapid Hypercasual Prototyping: Testing 10+ WebGL mechanics per day, flipping fully functional browser games on itch io or CodeCanyon, and licensing prototypes directly to casual game portals. 3. Interactive WebGL Client Solutions: Delivering custom 3D promotional browser games and interactive brand experiences for clients in 48 hours instead of weeks. First month results: > WebGL games generated: 24 > Viral impressions generated: 3.8M+ > Total revenue across licensing & content: $21,400 The AI completely automated the core development lifecycle: Claude Opus 5 built the physics engine, rendered 3D graphics in WebGL, hooked up audio controllers, and generated interactive browser logic with zero manual line-by-line coding. Bookmark it and check article 👇show more

Ridark
11,592 次观看 • 1 个月前