We are presenting **SAD** (Segment Any RGBD): SAD is... able to perform 3D segmentation (segment out any 3D object) with RGBD inputs (or rendered depth images only). - Code: - Demo Hugging Face:show more

Ziwei Liu
195,642 Aufrufe • vor 3 Jahren
Segment Any 3D Gaussians paper page: Interactive 3D segmentation... in radiance fields is an appealing task since its importance in 3D scene understanding and manipulation. However, existing methods face challenges in either achieving fine-grained, multi-granularity segmentation or contending with substantial computational overhead, inhibiting real-time interaction. In this paper, we introduce Segment Any 3D GAussians (SAGA), a novel 3D interactive segmentation approach that seamlessly blends a 2D segmentation foundation model with 3D Gaussian Splatting (3DGS), a recent breakthrough of radiance fields. SAGA efficiently embeds multi-granularity 2D segmentation results generated by the segmentation foundation model into 3D Gaussian point features through well-designed contrastive training. Evaluation on existing benchmarks demonstrates that SAGA can achieve competitive performance with state-of-the-art methods. Moreover, SAGA achieves multi-granularity segmentation and accommodates various prompts, including points, scribbles, and 2D masks. Notably, SAGA can finish the 3D segmentation within milliseconds, achieving nearly 1000x acceleration compared to previous SOTA.show more

AK
69,542 Aufrufe • vor 2 Jahren
Today we're releasing the Segment Anything Model (SAM) —... a step toward the first foundation model for image segmentation. SAM is capable of one-click segmentation of any object from any photo or video + zero-shot transfer to other segmentation tasks ➡️show more

AI at Meta
3,570,711 Aufrufe • vor 3 Jahren
NVIDIA finally released Neuralangelo's source code! The model can... turn videos from any device into detailed 3D structures, fully replicating buildings, sculptures, or other real aworld objects or spaces virtually. Here's how it works: A model utilizes a 2D video with multiple angles of an object or scene. I selects frames from different viewpoints to understand depth, size, and shape. The AI creates an initial 3D representation, similar to a sculptor shaping a subject. The render is optimized to enhance details, like a sculptor refining texture. The outcome is a 3D object or scene suitable for virtual reality, digital twins, or robotics.show more

Lior Alexander
478,039 Aufrufe • vor 3 Jahren
🚀 The Segment Anything Model (SAM) has been upgraded... to SAM2, featuring an efficient image encoder for segmenting images and videos. But does SAM2 outperform SAM1 in medical image and video segmentation? We're thrilled to present our paper "Segment Anything in Medical Images and Videos: Benchmark and Deployment"! We comprehensively benchmark SAM2 across 11 medical image modalities and videos. 📄 Paper: 💻 Code: **Highlights:** 1. SAM2 doesn’t always outperform SAM1 in 2D medical images, but excels in video segmentation, making it more accurate and efficient for 3D images, such as CT and MR scans. 2. MedSAM still outperforms SAM2 on most 2D modalities, but SAM2 surpasses MedSAM for 3D image segmentation in a slice-by-slice approach. 3. Segmentation performance varies with model size; sometimes the smallest model outperforms larger ones. 4. Fine-tuning SAM2 significantly boosts its performance for medical image segmentation. While SAM2 may struggle with challenging objects that have unclear boundaries or low contrast, it excels in generating good initial segmentation masks for common medical images and videos. However, the official interface doesn’t support medical data formats and has limitations on video length. To address this, we've developed a 3D Slicer Plugin and Gradio API for efficient 3D medical image and video segmentation. We invite you to try them out and provide feedback! 🔧 Deployment: - 3D Slicer Plugin: - Gradio API: (Note: Due to GPU limitations, the online API is available for only 12 hours and may be slow. We highly recommend deploying the Gradio API with your own computing resources: A big shoutout to Jun Ma (JunMa) who recently joined our UHN AI hub (UHN AI Hub) as Machine Learning Lead, and kudos to all co-authors: Sumin Kim, Feifei Li, Mohammed Baharoon (Mohammed Baharoon), Reza Asakereh, and Hongwei Lyu! This is true teamwork! Looking forward to collaborating with the community to advance 3D medical image and video segmentation foundation models! University Health Network U of T Department of Computer Science Department of Laboratory Medicine & Pathobiology Temerty Centre for AI in Medicine (T-CAIREM) Vector Institute #MedTech #AIinHealthcare #DeepLearning #MedicalImaging #SAM2 #MedSAM #AIResearchshow more

Bo Wang
178,539 Aufrufe • vor 2 Jahren
First Sachi 3D file is ready which you are... able to use your character where you want in any metaverse, for rendering or in animation as a few example. Use the link below to try it out👇show more

LUX
10,775 Aufrufe • vor 3 Jahren
Some of us come from humble families, old houses,... small towns with sad stories. We are not hustling to impress or be in any competition with anyone. We just want to change the storyline and fight the battles our parents never won. May we not labour in vain!show more

NaijaFarmer
19,514 Aufrufe • vor 1 Jahr
HTML enters 3D! Or vice versa? With the new... HTML in Canvas by WICG, we can finally put native DOM elements directly into WebGL/WebGPU scenes. It is experimental for now, but the possibilities for 3D interfaces and special effects are huge. This demo was built using Three.js and Omma AI (tool by Spline ) It’s a fun new way to explore what the web can do! Are you interested in seeing the demo?show more

Gábor Pribék
176,065 Aufrufe • vor 3 Monaten
I built a body for my openclaw I told... my agent that you have a ESP32-S3, 4 buttons and 3d printed body, and it looks like a milady. My , downloaded the eyes, brows and mouth, and configured its mood. Now it goes sad every one hour if I don't pet it or feed it. We have 3D printed 50 devices+ configured ESP32s and are looking to co build with devs around the world. If you are interested, drop a comment and text me in DMs with your address, will FedEx your way miladyshow more

Sarv
38,746 Aufrufe • vor 4 Monaten
messy inputs 👉🏽 polished outputs built a prototype with... the idea of letting users bring together images from Lummi into a canvas, quickly wireframe a concept, and then generate polished images as a photo, illustration, 3d render, whatever—with the right lighting, shadows, cohesive colors, and all that good stuff the output is still not great... but this is where I see the future of creative tools heading—kind of like how you give ChatGPT an idea for an email—all the messy bits—and it writes it for you in any tone or style you want in a clear way. now imagine that but for design or imageryshow more

Pablo Stanley
12,109 Aufrufe • vor 1 Jahr
2023 was the year of AI avatars 2024 was... the year of AI photos 2025 was the year of AI videos And I think it's becoming clear now that 2026 will be the year of AI world models Fully interactive explorable 3d worlds generated from one or multiple 2d images or a prompt In turn these 2d images can then be generated by AI too So soon you can generate fully explorable virtual 3d worlds based on your own imagination Next will be figuring out how to make those worlds interactive This is World Labs (unaffiliated, but I like it) As always a lot of big AI model companies are now working on the same thing: 3d world models, only World Labs has a real properly working demo (for now) Very exciting time again!show more

@levelsio
584,228 Aufrufe • vor 10 Monaten
It is sad to say but it looks like... I won’t be able to walk onto the ASU football team and will be closing that chapter. That said, I’m putting my name out there for any college coaches that are looking for a Kicker/Punter/Place Holder. •Current Junior 2 years eligibility •Planning to pursue masters •Never been on NCAA roster NCAA id & Hudl link are in my bio, film is on my page as well. If you are interested or need more info DMs are open. Still chasing my dream of playing college football 🙏 AlexZendejasKicking&Training Rausch Kicking Jaden Oberkromshow more

Nolan Krinsky
28,154 Aufrufe • vor 1 Jahr
DimensionX: Create Any 3D and 4D Scenes from a... Single Image with Controllable Video Diffusion TL;DR: Create 3/4DGS from Video Diffusion Note: Some first inference code released (not all yet). Contributions (cited): • We present DimensionX, a novel framework for generating photorealistic 3D and 4D scenes from only a single image using controllable video diffusion. • We propose ST-Director, which decouples the spatial and temporal priors in video diffusion models by learning (spatial and temporal) dimension-aware modules with our curated datasets. We further enhance the hybriddimension control with a training-free composition approach according to the essence of video diffusion denoising process. • To bridge the gap between video diffusion and real-world scenes, we design a trajectory-aware mechanism for 3D generation and an identity-preserving denoising approach for 4D generation, enabling more realistic and controllable scene synthesis. • Extensive experiments manifest that our DimensionX delivers superior performance in video, 3D, and 4D generation compared with baseline methods.show more

MrNeRF
17,052 Aufrufe • vor 1 Jahr
✨ Made a new mini feature on Photo AI:... [ Grab from 3d model ] So the problem is we're at that stage in time (typical for AI) where image-to-3d models are not good enough but are fun to play with, but we know they'll be good enough in 1-2 years With [ Make 3d model ] you already can turn any Photo AI pic into a 3d model but it still looks hyper clunky and deformed, but it works! One cool idea I had to make that more useful and made now: Let people make a 3d model then change the view of the it with the 3d viewer, then press [ o ] and it grabs a frame of the 3d That image you can then [ Remix ] (img2img), and it becomes a real photo again and that in turn you can then turn into a video again with [ Make video ] So that essentially gives you a fully freeform camera position control to take photos with One thing I need to fix is the background/skybox, I kinda need to take the original photo and remove the person and just get the background for the 3d model viewer, in this case it should be white, but it's a start!show more

@levelsio
119,210 Aufrufe • vor 1 Jahr
this is the moment successfully finished scheduling shifts for... my parent's tea store for the first time. my mom is BLOWN AWAY. this is going to save her hours of every week going back and forth with the team and sorting out this annoying task. the way we set it up: - "CameliaOS", our moltbot, sends her a reminder every morning to ask for inputs from team - team members respond with times - my mom sends screenshots to Camelia - Camelia updates her on any missing inputs - Camelia drafts a plan, adds it to google calendar (color coded, which is the form factor that the team is used to) - my mom gives Camelia feedback - she ships it one interesting learning: as we were teaching Camelia this new skill, my mom sometimes asked me "how should i explain to it that X or Y" I told her - just discuss with her like you would with any human, just share your doubts and concerns.. Opus is so great at making sense from vague inputs and help crystalize a plan. we are making progress, many more workflows to teach CamealiaOS, and I am sure more will surface the deeper we go into it.show more

Dan Peguine
266,374 Aufrufe • vor 5 Monaten
Apple just trained a 3D Gaussian head reconstruction model... on 10,000+ subjects. Feed-forward. No test-time optimization. New identity in, reconstructed Gaussian head out. The UV-parameterized Gaussian representation decouples the number of Gaussians from the number and resolution of input images, making it practical to train with many high resolution views. And the heads are not just static either: text-conditioned identity generation, plus blendshape-driven latent animation across identities. We've been building in the 3D Gaussian Splatting space for a while. The gap between "research demo" and "works on real people at scale" is closing fast.show more

KIRI Engine - 3D Scanner App
12,181 Aufrufe • vor 2 Monaten
Depth Any Video with Scalable Synthetic Data AI physicists... and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.show more

MrNeRF
27,428 Aufrufe • vor 1 Jahr
Teaching robots how to paint! 💅🏼 This painting robot,... for example, mimics human movements with precision to handle repetitive, and tiring tasks. A special device memorizes points in 3D space, which are then sent to the robot's control. The result? A robot that, after a single demonstration, can perform a given action. Keep in mind that it's taught as 'fixed'. 👨🏻🔧 That is, it will not be able to react to an anomaly or changing environment. But ultimately, with fewer workers available and many avoiding tough, dirty jobs, robots are helping industries stay efficient and safe. ♻️ RT to help 1 robot find a new workplace!show more

Lukas Ziegler
38,319 Aufrufe • vor 1 Jahr