正在加载视频...

视频加载失败

HunyuanWorld-Voyager is here and fully open-source! The world’s first ultra-long-range world model with native 3D reconstruction, redefining AI-driven spatial intelligence for VR, gaming, and simulations. ✅Direct 3D Output: Exports point cloud videos to 3D formats without tools like COLMAP, enabling instant 3D application use. ✅Innovative 3D Memory: Introduces a...

198,289 次观看 • 11 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

World Model is trending— let's revisit our HunyuanWorld journey. We’ve been pioneering open-source 3D world generation in the past two months, and this ride’s only getting started. 🌍 📅 July: HunyuanWorld 1.0 📌 First open-source 3D world model compatible with CG pipelines (Unity/Unreal/Blender) 📌 Hit 2K+ GitHub stars in just two months ⭐—thank you for the love! 📅 August: 1.0-Lite 📌Same top-tier quality, running on consumer GPUs! 📅 September: 1.0-Voyager 📌 Direct 3D output + world memory—taking exploration further! Seamlessly integrated into CG pipelines with layered 3D modeling (assets, terrain, skybox) and fully open-sourced.. we’re fully committed to building open-source spatial intelligence for all! 🚀 💡 Why it matters? ✅ Seamless CG Pipeline Integration: Export generated 3D scenes as standard mesh formats, effortlessly integrating into industry-standard tools like Blender, Unity, and Unreal Engine for direct editing, animation, and physical simulation. ✅ Hierarchical Scene Editing: Deconstruct scenes into semantic layers (sky, background, foreground objects) via instance recognition and layer decomposition, allowing for atomic-level control—independently modify, relocate, or replace objects without rebuilding the entire world. Project page: Github: Amazing creations by Stijn Spanhove camenduru GENEL | AIを用いた動画制作 apolinario 🌐 とりにく Directive Creator 🪥 👇 #AI #3DGeneration #OpenSource #WorldModels #Hunyuan3D #HunyuanWorld

Tencent HY

20,178 次观看 • 10 个月前

🚀 Announcing Echo — our new frontier model for 3D world generation. Echo turns a simple text prompt or image into a fully explorable, 3D-consistent world. Instead of disconnected views, the result is a single, coherent spatial representation you can move through freely. This is part of a bigger shift in AI: from generating pixels and tokens to generating spaces. Echo predicts a geometry-grounded 3D scene at metric scale, meaning every novel view, depth map, and interaction comes from the same underlying world — not independent hallucinations. Once generated, the world is interactive in real time. You control the camera, explore from any angle, and render instantly — even on low-end hardware, directly in the browser. High-quality 3D world exploration is no longer gated by expensive equipment. Under the hood, Echo infers a physically grounded 3D representation and converts it into a renderable format. For our web demo, we use 3D Gaussian Splatting (3DGS) for fast, GPU-friendly rendering — but the representation itself is flexible and can be easily adapted. Why this matters: consistent 3D worlds unlock real workflows — digital twins, 3D design, game environments, robotics simulation, and more. From a single photo or a line of text, Echo builds worlds that are reliable, editable, and spatially faithful. Echo also enables scene editing and restyling. Change materials, remove or add objects, explore design variations — all while preserving global 3D consistency. Editing no longer breaks the world. This is only the beginning. Echo is the foundation for future world models with dynamics, physical reasoning, and richer interaction — environments that don’t just look right, but behave right. Explore the generated worlds on our website and sign up for the closed beta. The era of spatial intelligence starts here. 🌍 #Echo #WorldModels #SpatialAI #3DFoundationModels Check it out:

SpAItial AI

176,105 次观看 • 7 个月前

Wow. Recreating the Shawshank Redemption prison in 3D from a single video, in real time (!) Just read the MASt3R-SLAM paper and it's pretty neat. These folks basically built a real-time dense SLAM system on top of MASt3R, which is a transformer-based neural network that can do 3d reconstruction and localization from uncalibrated image pairs. The cool part is they don't need a fixed camera model -- it just works with arbitrary cameras -- think different focal lengths, sensor sizes, even handling zooming in video (FMV drone video anyone?!). If you've done photogrammetry or played with NeRFs you know that is a HUGE deal. They've solved some tricky problems like efficient point matching and tracking, plus they've figured out how to fuse point clouds and handle loop closures in real-time. Their system runs at about 15 FPS on a 4090 and produces both camera poses and dense geometry. When they know the camera calibration, they get SOTA results across several benchmarks, but even without calibration, they still perform well. What's interesting is the approach -- most recent SLAM work has built on DROID-SLAM's architecture, but these folks went a different direction by leveraging a strong 3D reconstruction prior. Seems to give them more coherent geometry, which makes sense since that's what MASt3R was designed for. For anyone who cares about monocular SLAM and 3D reconstruction, this feels like a significant step toward plug-and-play dense SLAM without calibration headaches -- perfect for drones, robots, AR/VR -- the works!

Bilawal Sidhu

703,928 次观看 • 1 年前

🚨 SIGGRAPH Asia 2025 Paper Alert 🚨 ➡️Paper Title: WorldExplorer: Towards Generating Fully Navigable 3D Scenes 🌟Few pointers from the paper 🎯Generating 3D worlds from text is a highly anticipated goal in computer vision. Existing works are limited by the degree of exploration they allow inside of a scene, i.e., produce stretched-out and noisy artifacts when moving beyond central or panoramic perspectives. 🎯 To this end, authors of this paper proposed “WorldExplorer”, a novel method based on autoregressive video trajectory generation, which builds fully navigable 3D scenes with consistent visual quality across a wide range of viewpoints. 🎯They initialize their scenes by creating multi-view consistent images corresponding to a 360 degree panorama. 🎯Then, they expanded it by leveraging video diffusion models in an iterative scene generation pipeline. 🎯Concretely, they generated multiple videos along short, pre-defined trajectories, that explore the scene in depth, including motion around objects. 🎯Their novel scene memory conditions each video on the most relevant prior views, while a collision-detection mechanism prevents degenerate results, like moving into objects. 🎯Finally,they fuse all generated views into a unified 3D representation via 3D Gaussian Splatting optimization. 🎯Compared to prior approaches, WorldExplorer produces high-quality scenes that remain stable under large camera motion, enabling for the first time realistic and unrestricted exploration. 🎯They believe this marks a significant step toward generating immersive and truly explorable virtual 3D environments. 🏢Organization: TU München 🧙Paper Authors: Manuel-Andreas Schneider, Lukas Höllein , Matthias Niessner 📝 Read the Full Paper here: 🗂️ Project Page: 🧑‍💻 Code: 🎥 Be sure to watch the attached Technical Summary Video - Sound on 🔊🔊 Find this Valuable 💎 ? ♻️QT and teach your network something new Follow me 👣, naveen manwani , for the latest updates on Tech and AI-related news, insightful research papers, and exciting announcements. #SIGGRAPHAsia2025

naveen manwani

10,578 次观看 • 10 个月前

Here are 10 AI video editor GitHub repos worth bookmarking: 1. Shotcut Most actively maintained open source video editor in 2026. 14K stars. Cross-platform with AI-assisted features. Just shipped a new release April 30, 2026. 2. Kdenlive The closest open source alternative to Adobe Premiere Pro. Multi-track editing, proxy editing, VST audio, and customizable workspace. Best for professional workflows. 3. OpenShot The easiest entry point for beginners. Drag and drop, 400+ transitions, 3D titles, and AI-assisted trimming. 5,700 stars. 4. Blender Not just 3D. Blender's video sequence editor and compositing pipeline is used in professional film production. 18,300 stars. Unmatched for VFX. 5. Recordly Screen recorder with auto-zoom, cursor polish, webcam overlays, and styled frames built in. Built for demo videos and walkthroughs. 6. Wan2.1 Alibaba's open source text-to-video model. Cinema-grade 1080p generation. Apache 2.0. The gold standard for open source video generation in 2026. 7. HunyuanVideo Tencent's 13B parameter open source video model. 11.9K stars. Handles 720p and 1080p with high temporal coherence. 8. CogVideoX Apache 2.0 licensed. Loads natively via Hugging Face Diffusers. Strong prompt following and smooth frame transitions. Needs 16GB VRAM minimum. 12.5K stars. 9. Open-Sora Most starred open source video generation project at 24K stars. Full training pipeline for $200K. Production-level output quality. 10. Mochi 1 Focused entirely on motion quality. The most natural-looking physics of any open source video model. Water, fabric, and human gestures without AI jitter. Apache 2.0.

Kanika

17,726 次观看 • 1 个月前

🚀 The Future of 3D Creation is Here! World Labs World Labs dropped a bombshell with #Marble, and the spatial computing world is buzzing. 🤯 👉 For Beginners: Create stunning 3D worlds, no code needed. 👉 For Pros: Power-up your workflow with advanced tools. 👉 For Industry: Marks a milestone in Spatial AI innovation! But let's talk about the elephant in the room: Hardware. 🖥️ While #Marble leverages cloud power to generate worlds in seconds, accessing and sharing them still depends heavily on local device performance. 𝗧𝗵𝗲 𝗖𝗵𝗮𝗹𝗹𝗲𝗻𝗴𝗲𝘀 𝗔𝗿𝗲 𝗥𝗲𝗮𝗹: ➤ 𝗔𝗰𝗰𝗲𝘀𝘀𝗶𝗯𝗶𝗹𝗶𝘁𝘆: Basic laptops struggle to run initial generations smoothly — not to mention high-fidelity scenes. Sharable links are a step forward, but must every viewer own high-end hardware? ➤ 𝗠𝗼𝗯𝗶𝗹𝗲 𝗚𝗮𝗽: Full creation tools aren’t mobile-ready yet. Desktop is still a must for the complete experience. ➤ 𝗩𝗥 "𝗩𝗶𝗲𝘄𝗶𝗻𝗴": "Open in VR" is a great start—but true collaborative VR requires serious rendering power and device support. How do we include everyone, without top-tier headsets or perfect networks? We tested #Marble on a typical Windows laptop. See the performance gap? We bring fluid Marble experience via #LarkXR! 👀 ▶︎ 𝗙𝗹𝗶𝗽 𝘁𝗵𝗲 𝗦𝘄𝗶𝘁𝗰𝗵 #LarkXR is here. Our Cloud XR Streaming Technology delivers high-fidelity VR and 3D experiences to any device — breaking through hardware limits and compatibility gaps. Now every user can join, view, and collaborate seamlessly. #Marble Unlocks Spatial World Creation. #LarkXR Unlocks Interactive Immersive Accessibility. 𝘓𝘦𝘵'𝘴 𝘸𝘰𝘳𝘬 𝘵𝘰𝘨𝘦𝘵𝘩𝘦𝘳 𝘵𝘰 𝘦𝘯𝘴𝘶𝘳𝘦 𝘦𝘷𝘦𝘳𝘺𝘰𝘯𝘦, 𝘦𝘷𝘦𝘳𝘺𝘸𝘩𝘦𝘳𝘦, 𝘰𝘯 𝘢𝘯𝘺 𝘥𝘦𝘷𝘪𝘤𝘦, 𝘤𝘢𝘯 𝘵𝘳𝘶𝘭𝘺 𝘴𝘵𝘦𝘱 𝘪𝘯𝘵𝘰 𝘵𝘩𝘦 𝘸𝘰𝘳𝘭𝘥𝘴 𝘺𝘰𝘶 𝘤𝘳𝘦𝘢𝘵𝘦. 𝗚𝗲𝘁 𝗦𝘁𝗮𝗿𝘁𝗲𝗱 𝗪𝗶𝘁𝗵 𝗟𝗮𝗿𝗸𝗫𝗥 𝗧𝗼𝗱𝗮𝘆 🚀 🚀 🚀 #WorldLabs

Paraverse

16,000 次观看 • 8 个月前

GeoLibre v2.0.0 is here! GeoLibre is a free and open-source geospatial platform that runs everywhere: as a native desktop app, in the browser, on Android, and embedded right inside Jupyter notebooks. It brings modern web mapping, cloud-native data formats, and a full processing toolbox together in one place, all built on MapLibre and with no proprietary lock-in. Our first major release adds a true 3D globe, takes mapping beyond Earth to Mars and the Moon, lets styles round-trip with QGIS, and turns loaded vector layers into editable, save-back-to-source data. What's new in v2.0.0 - Planetary mapping: explore Mars, the Moon, and other bodies with the OpenPlanetaryMap basemaps, a per-project ellipsoid, and a planet switcher right in the Layers panel. - CesiumJS 3D globe: switch any map pane to a photorealistic 3D globe that stays camera-synced with your 2D maps and mirrors the layer stack. - True 3D data: render vector layers with Z coordinates, load TIN/MultiPatch 3D shapefiles, and display KML/KMZ Collada (.dae) 3D models. - Symbology interchange: import and export vector styling as OGC SLD, QGIS QML, and Mapbox GL style JSON, so styles round-trip between GeoLibre, QGIS, and the Mapbox/MapLibre ecosystem. - Editable source layers: edit vector layers and write the changes back to their source, including GeoPackage and GeoJSON files and PostGIS database tables. - Weather and sky: a new Weather menu with live cloud and precipitation radar overlays (RainViewer), plus a Google Earth-style sun position simulation for realistic lighting. - Terrain and lighting: double-click the terrain control to set vertical exaggeration, and view any scene in true 3D relief. - Smarter data import: bring in CSV without coordinates as an attribute table, split GPX track points and route points into separate layers, and load macOS-zipped and projected-CRS shapefiles. - Raster in the browser: build normalized-difference indices for any HTTP COG and extract COG/WMS/XYZ bounding-box subsets client-side. - Field Calculator upgrades: compute geometry length and area directly on your features. - Attribute table: multi-select rows with Ctrl and Shift, plus faster navigation. - Google Earth-style extras: "View in Google Maps / Google Earth" actions, camera-reset keyboard shortcuts, and a UTM easting/northing grid mode for the Gridlines overlay. - New plugins: a Mapillary coverage and street-level image viewer, a Historical Imagery panel, and an Elevation Profile tool. - Fully localized: all 13 language catalogs are complete, so the entire UI is translatable. Try it out - Launch GeoLibre Web: - GitHub: - Documentation: - Release notes: #GIS #GeospatialData #OpenSource #RemoteSensing #DataVisualization #MapLibre #GeoLibre

Qiusheng Wu

35,499 次观看 • 20 天前

🚀 We’re hiring! Staff Scientist / Postdoc – Tissue Clearing & 3D Image Analysis (m/f/d) (LMU Munich) Are you a great fit, or do you know someone outstanding, please reach out 🔁 If you want to at the frontier of whole-organ / whole-body 3D imaging, and help generate truly beautiful datasets that drive major biological discoveries and therapeutic development, see below ✨ We’re building the next-generation pipeline for tissue clearing + light-sheet microscopy + quantitative 3D analysis in the SyNergy Excellence Cluster (Mesoscale Hub) and we’re looking for someone excited to push this forward with us. 🧠🔬📈 🎥 I’m also attaching a short video showing the kind of high-quality imaging and datasets you’d be working with. What you’ll do 🛠️ 🔹 Lead and evolve tissue clearing + light-sheet workflows across collaborative SyNergy projects 🔹 Turn complex 3D datasets into robust quantitative insights (visualization, atlas registration, readouts) 🔹 Develop new methods and analysis pipelines together with our AI team 🤖 🔹 Maintain and optimize cutting-edge light-sheet systems (optional: support animal license writing) What we’re looking for 🎯 ✅ Strong hands-on experience in tissue clearing and/or fluorescence microscopy ✅ Solid experience with light-sheet microscopy and 3D imaging workflows ✅ Familiarity with 3D tools like Imaris / arivis Vision4D, stitching (e.g., BigStitcher), and quantitative analysis in cleared tissues ✅ Service mindset, great organization, and strong scientific English How to apply 📩 Apply via the LMU Klinikum online application form Please also send your application to: [email protected] CC: [email protected] 📎 Include one PDF: short cover letter, CV, 2–3 referees, and earliest start date. 📍 Campus Großhadern (Munich) and Helmholtz Munich | 🕒 Full-time | 📅 Start: 01 January 2026 If you love high-quality imaging, cutting-edge biology, and building something that will matter, we’d love to hear from you. 🌍✨ #hiring #StaffScientist #Postdoc #TissueClearing #LightSheetMicroscopy #ImageAnalysis #SpatialBiology #Neuroscience #SyNergy #LMU #Munich

Ali Max Erturk

14,751 次观看 • 7 个月前

🔴 Finally! NVIDIA has finally made the code for Neuralangelo public! It has the ability to transform any video into a highly detailed 3D environment, and it's a technology related to but DIFFERENT from NeRF. 💡 Here's how it works: It takes a 2D video as input, showing an object, monument, building, landscape, etc., from various perspectives and analyzes details such as depth, size, and the shapes of objects. From this, the AI sketches an initial 3D model, similar to how an artist molds a figure. This representation is then refined to highlight more details, just as an artist would make the final touches when sculpting. The result is a 3D environment/model, perfect for use in any environment. Imagine the applications it will have for video games, cinema, virtual environments, VR, and more! 📽️🎮 💡 More details: A year ago, an article was presented on a groundbreaking technique called NVIDIA's Instant NeRF. This technique turns images into stunning 3D scenes in a short time, ideal for creating realistic models for video games and other applications. Although Instant NeRF had a lot of potential, the generated models were not perfect and often lacked detailed structures, appearing somewhat cartoonish. A year on, NVIDIA releases a new technique based on Instant NeRF, named Neuralangelo. This enhances the fidelity of surface structures. While NeRF reconstructs real objects in virtual environments from images or videos, Instant NeRF speeds up this process, and Neuralangelo further improves the quality, making the generated objects appear even more realistic when examined up close. Neuralangelo improves Instant NeRF's approach in two key ways related to the hash grid encoding technique: 1⃣ Numerical gradients have been used to compute higher-order derivatives as a smoothing operation. This optimizes the "hash grid" encoding using numerical rather than analytical gradients, providing a smoother input to the network that produces the 3D model. 2⃣ A "coarse-to-fine" optimization has been implemented in the hash grids to control different levels of detail. That is, they first focus on a smoothed version of the scene, and then refine it with more detailed updates. Well, as Arthur C. Clarke said, "Any sufficiently advanced technology is indistinguishable from magic."

Javi Lopez ⛩️

689,169 次观看 • 3 年前