Bai et al., "Positional Encoding Field" Make your RoPE... encoding 3D by including a z axis, then manipulate your image by simply manipulating your positional encoding in 3D --> novel view synthesis. Neat idea.show more

Kwang Moo Yi
46,880 views • 9 months ago
one 2d photo --> 3d gaussian splat quick test... with Echo-2 by SpAItial -- these 3d scene generation models are getting better! already at a sufficient quality to serve as a virtual set / backdrop in your 3d tool of choiceshow more

Bilawal Sidhu
23,511 views • 2 months ago
Wonderland: Navigating 3D Scenes from a Single Image Contributions:... • First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. • Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. • Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasks—and bridging image space and 3D space—through the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256× spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.show more

MrNeRF
52,801 views • 1 year ago
Combining the explicit control of 3D software with the... creativity of generative AI models is a promising yet underrated workflow. Build your 3D scenes procedurally by describing them in natural language, then take them all the way with your image & video models of choice. Tools like intangible are built around such a workflow so you don't need to duct-tape apps together. Pretty cool!show more

Bilawal Sidhu
37,629 views • 1 year ago
simple character design workflow in Freepik spaces, with total... control over your creations > create the character using NB Pro nodes > generate 3D views > integrate it into a midjourney environment > animate on Kling 2.6 character design inspired by Scopper Gabanshow more

INK
49,784 views • 7 months ago
✨ 3d models are now LIVE on Photo AI... 😊 You can now turn any AI photo you make into a 3d model by pressing [ 📦 Make 3d model ] And then you can view it inside Photo AI or download it as a .GLB 3d model file It's still very early in AI generated 3d model world but it's nice to have this feature working already As always, the models will keep improving, so this feature will keep getting better (like it did with video, it sucked before, now it's getting passable) Next would be nice to switch to .USDZ so you can load it straight into your iPhone with ARKit and put it in your room Available now for everyone on the Premium and Ultra planshow more

@levelsio
112,930 views • 1 year ago
📢Announcing our 3D head avatar benchmark📢 Two tasks with... hidden test sets: - Dynamic Novel View Synthesis on Heads - Monocular FLAME-driven Head Avatar Reconstruction Our goal is to make research on 3D head avatars more comparable and ultimately increase the realism of digital humans. The benchmark studies distinct phenomena of 3D head avatar creation, such as extreme facial expressions, slow motion captures of shaking long hair, or complicated light reflection and refraction patterns of glasses. The two benchmark tasks assess two core desiderata of 3D avatars: While the novel view synthesis challenge focuses on best possible rendering quality of complex moving scenes, the avatar animation challenge is concerned with how well a driving signal is translated into an avatar. Evaluations are light-weight and consist of diverse video recordings from the popular NeRSemble dataset with a hidden test set. Participation in the benchmark is therefore straight-forward and requires only 5 reconstructions per task. Leaderboard and benchmark submission: Benchmark data access and toolkit: Great work by Tobias Kirschstein Simon Giebenhainshow more

Matthias Niessner
28,075 views • 1 year ago
When parking, you can now see a high fidelity... 3D representation of the world around your vehicle, including proximity & shape of nearby objects, barriers, vehicles & painted road markings By using a dedicated neural network to model obstacles & paint lines, we can accurately estimate distances & represent arbitrary shapes in a smooth & computationally efficient wayshow more

Tesla
1,550,361 views • 2 years ago
If you typically stream with a 3D model, I... highly recommend that you pose your model when you’re not doing full body mocap instead of letting it stay in the stiff generic pose! It makes a huge difference in how your energy is conveyed (ᗒ⩊ᗕ)⸝ި ʕᦏ⌎ I’m only saying this because I noticed this a lot with my own 3D kids, but here’s a comparison showcase: ← default pose & default pendulum physics in Warudo → custom poses & slightly adjusted pendulum physics If you find it hard to pose within Warudo by using bone offsets or none of the existing poses in Warudo vibe with you, you can make your own like me! I made my custom poses in “VRM Posing Desktop” (this is the BEST vrm posing app I’ve ever used in the past 3 years) and exported them as Unity anims then dropped them into Warudo’s animation folder! You can then make a simple blueprint in Warudo to toggle between poses and make your model look more alive! It should fit especially well for just chatting streams 🙂↕️✨show more

𝗞𝗔𝗥𝗜𝗛𝗔 🌘🍀 3D Artist ☻
15,700 views • 4 months ago
How to generate 3D miniature city models and animate... them using Kling AI? This visual effect can be created with Kling O1, and then rotated in 3D using Image to Video. The image prompt used for generation is as follows: Present a clear, 45° top-down isometric miniature 3D cartoon scene of New York featuring its most iconic landmarks and architectural elements. Use soft, refined textures with realistic PBR materials and gentle, lifelike lighting and shadows. Integrate the current weather conditions directly into the city environment to create an immersive atmospheric mood.Use a clean, minimalistic composition with a soft, solid-colored background. At the top-center, place the title "New York" in large bold text in white. You can create different effects by changing the city name according to your needs.show more

Kling AI
34,303 views • 7 months ago
Everyone's sleeping on image-to-3D AI models. They can make... your app look incredibly unique, with just a little effort. Here's how. This is my calorie tracker, built in a week with nothing but prompting. Just Claude Code + a couple APIs. The visuals are all AI-generated. I'll be sharing the full workflow + all the crazy technical stuff Claude and I did to make this work, so nobody has to struggle through it like me. Deep dive coming soon! Till then, this is the high-level idea: 1. Get a clean image of the food (or whatever your asset is) - In my app, the user describes foods via text, or attaches images (or both) - If text, an LLM extracts the food description and formats it into a specific prompt I tuned for this design, and we generate an image using Z-Image Turbo through fal - If image, we do the same thing but with FLUX.2 [dev] to edit the user image into our reference design - Originally, both used Google Nano Banana, but switching to open models cut costs and latency a ton 2. Gaussian splatting (2D image → 3D model) - I tried various 2D-to-3D options on fal and ended up with TripoSplat as my preferred balance of speed, cost, latency; this turns an image into a 3D model that looks super high quality (link below) - The app displays the 2D image while our backend generates the 3D splat - We "groom" the splat to reduce size and load time by culling low-opacity/scale points 3. Render efficiently on device Originally, it looked great but ran at 10 FPS. Getting to 120 FPS was a crazy journey. TL;DR: - SwiftUI had to go; it forced us to render each asset in independent MTKViews, which wasn't workable - Instead, we composite every dish into one full-bleed CAMetalLayer using MetalSplatter (link below) - We had to make some optimizations within MetalSplatter's code too, to reduce the overhead of sorting points per render Then I added some finishing touches like the subtle rotation and parallax as they move around. I think it turned out pretty cool :) Overall, this took some effort, but we still got it done in less than a day. Hopefully your agent can follow in the footsteps of mine and do it much faster. Keep an eye out for the bigger writeup, which'll give your agent everything it needs. If you have any questions, drop em below!show more

Anshu
19,931 views • 1 month ago
GLORIETTA'S LED AD FOR HAN IS NOW DISPLAYED! Come... and get by at Glorietta Activity Center to view this 3D Birthday LED Ad for Hani. 📌 For those our HAN925FM cupsleeve attendees who opted for a meetup, you may now claim your kits from our admins in front of the LED ad. Thank you! FEELING 22 WITH HANI #HANforgettableDay #달콤한이_데이 #HappyHanDayshow more

PARK HAN PH
13,005 views • 10 months ago
Introducing Kaleido💮 from AI at Meta — a universal... generative neural rendering engine for photorealistic, unified object and scene view synthesis. Kaleido is built on a simple but powerful design philosophy: 3D perception is a form of visual common sense. Following this idea, we formulate rendering purely as a sequence-to-sequence generation problem, successfully unifying neural rendering with the architecture principles behind modern language and video models. Unlike traditional neural rendering methods, Kaleido learns 3D purely in a data-driven way, without explicit 3D representations or structures. It acquires spatial understanding directly through large-scale video pretraining, then multi-view 3D data finetuning, inspired by how LLMs acquire textual common sense from large corpora before specialising in domains like coding. Through extensive ablations, we progressively modernised the architecture design and training strategies and tackled key scaling challenges in sequence-to-sequence generative rendering, arriving at a design that’s simple, versatile, and scalable. Kaleido significantly outperforms prior generative models in few-view settings, and remarkably is the first zero-shot generative method matches InstantNGP-level rendering quality in multi-view settings. We view Kaleido also as an alternative step towards world modeling that flexibly spans a spectrum of “realities": with many views, it faithfully reconstructs grounded reality; with fewer views, it imagines plausible unseen details. 🔗 Explore more results and paper:show more

Shikun Liu
22,332 views • 9 months ago
🌍 As some of you might know, last year... we started building an app that required a 3D Map, and we were taken aback by the lack of good SDKs. They’re all clunky, slow, and unstable.🤔 Today, we're thrilled to introduce Cartes - a fast, easy-to-use, and visually appealing 3D Map SDK for Unity. This tool is built to empower your creativity in developing delightful apps and games for the Real World Metaverse. 📱🤳 ⚡️ With blazing-fast performance, our SDK offers a seamless integration for your (modern) Unity projects. We provide default navigation features, intuitive gestures, and clustering capabilities that we meticulously refined over hundreds of hours and proof-tested in guerilla tests. 🏞️ We wanted to build a genuinely 3D map, with 3D terrain, monuments and decorations, able to transition smoothly from a global view down to the human eye level. 🔍 #Unity #AR #Maps #LBEshow more

Tina Debove ᯅ
15,210 views • 3 years ago
✨ I can now generate 3d assets for my... drone sim at directly from Cursor (sponsor of #vibejam) I need buildings that you'd see in a war torn city, like warehouses in ruins, broken down abandoned houses, bombed out bridges etc. Nano Banana Pro or 2 can generate them really well and then you can put them in an image-to-3d model and you get a GLB or FBX That one you can then import into your Three.js game, the models might be big though, in my case like 16MB, so I ask it to compress it and make it more low poly so it loads fast ThreeJS then loads the individual GLBs on page load and puts them in my drone sim somewhere randomly, I think I should remove some of the grass and match the sandy color of the ruins though to make it fit in moreshow more

@levelsio
134,777 views • 3 months ago
2023 was the year of AI avatars 2024 was... the year of AI photos 2025 was the year of AI videos And I think it's becoming clear now that 2026 will be the year of AI world models Fully interactive explorable 3d worlds generated from one or multiple 2d images or a prompt In turn these 2d images can then be generated by AI too So soon you can generate fully explorable virtual 3d worlds based on your own imagination Next will be figuring out how to make those worlds interactive This is World Labs (unaffiliated, but I like it) As always a lot of big AI model companies are now working on the same thing: 3d world models, only World Labs has a real properly working demo (for now) Very exciting time again!show more

@levelsio
584,228 views • 10 months ago
Bug reporting is an essential part of making sure... Discord is up and running the way it's supposed to. Notice a glitch that's getting in the way of your chatting experience? Make sure to submit a bug report to our team by including: - A description of the bug - Steps to reproduce - Expected results - Actual results - Discord client info - Debug logs - An image/video attachment of what you're seeing Learn more about bug reporting:show more

Discord Support
36,945 views • 1 year ago
I created a dance movement sheet by using a... reference image to animate every 16 panels from the reference image I provided. GPT Image 2 + See Dance 2.0 on Yapper Tutorial Below Prompt Here’s every step with the text under each heading: 1. Basic Stance Stand with your feet shoulder-width apart. Relax your knees. Keep your upper body relaxed. Get ready. 2. Step to the Right Step to the right. Move your body in the direction of the step. Keep your knees soft. Keep your gaze forward. 3. Step to the Left Step to the left. Move your body in the direction of the step. Keep your knees soft. Return to the starting position. 4. Two Steps Take two steps in sequence. Step to the right first. Then step to the left. Connect smoothly. 5. Body Wave Start the wave from your chest. Move through your ribs and hips. Finish with your lower body. Make the motion smooth like a wave. 6. Hip Sway Move your hips side to side. Shift your weight with the motion. Use your body naturally. Keep your upper body relaxed. 7. Arm Swing Swing your arms wide. Step to beat with your feet. Coordinate your arms and steps. Return your arms to center. 8. Turn Preparation Prepare for a turn in place. Cross one foot over the other. Use your core to maintain balance. Spot in the direction of the turn. 9. Right Turn Turn to the right. Keep your core engaged. Pivot on the ball of your foot. Spot forward after the turn. 10. Left Turn Turn to the left. Keep your core engaged. Pivot quickly. Keep your posture upright. 11. Jump Up Bend your knees and jump up. Reach your arms overhead. Land softly on your feet. Connect to the next move. 12. Kick Pose Kick your leg forward. Keep your supporting leg stable. Engage your core for balance. Control your landing. 13. Side Lunge Step wide to one side. Bend one knee and lower your body. Keep the other leg straight. Show the direction of your strength. 14. Freeze Pose Hit a strong pose on the beat. Freeze your body momentarily. Create a powerful shape. Show your presence. 15. Finishing Pose Finish the dance gracefully. Keep your balance. Show your confidence. End with a clean, sharp pose. 16. Whole Flow Connect from the basic stance to the finishing pose. The steps, waves, and turns flow naturally together. Express your own style.show more

Sharon Riley
54,571 views • 3 months ago
So you want to become a VTuber in 2025?!... Here is a guide on how to get started, made by 4 year full-time VTuber and Manager, Lulu! 1. Know your goals Know if you are getting into VTubing as a hobby or as a business from the start and set your expectations accordingly 2. Start small F2U and P2U models and assets are available plenty via Nizima and Booth, VRoidstudio is free to make a 3D avatar with 3. Test the waters Make sure that streaming is possible for you, test your tech, your upload/download speeds, your voice, your set-up. Make sure that you give yourself a good experience as well as your future audience 4. Take your time Getting to affiliate/monetization will take time, getting used to the tech side of things takes time. It's unlikely you will wake up to 1k viewers overnight, and not every day will be great interactions only 5. Connect with other VTubers Make space and time for new friends, don't see them as networking opportunities or competition, but people on the same path. 6. Grow organically Using Ads or shifty tactics to stand out can not only jeopardize your standing in the community but can cost you your channel 7. Don't give up Even on days when things are difficult, don't give yourself too much pressure. Enjoy what you can and don't let other people or tech problems get you down, the journey is worth it.show more

Kuromiya Lucien
22,163 views • 1 year ago
🚀 The Segment Anything Model (SAM) has been upgraded... to SAM2, featuring an efficient image encoder for segmenting images and videos. But does SAM2 outperform SAM1 in medical image and video segmentation? We're thrilled to present our paper "Segment Anything in Medical Images and Videos: Benchmark and Deployment"! We comprehensively benchmark SAM2 across 11 medical image modalities and videos. 📄 Paper: 💻 Code: **Highlights:** 1. SAM2 doesn’t always outperform SAM1 in 2D medical images, but excels in video segmentation, making it more accurate and efficient for 3D images, such as CT and MR scans. 2. MedSAM still outperforms SAM2 on most 2D modalities, but SAM2 surpasses MedSAM for 3D image segmentation in a slice-by-slice approach. 3. Segmentation performance varies with model size; sometimes the smallest model outperforms larger ones. 4. Fine-tuning SAM2 significantly boosts its performance for medical image segmentation. While SAM2 may struggle with challenging objects that have unclear boundaries or low contrast, it excels in generating good initial segmentation masks for common medical images and videos. However, the official interface doesn’t support medical data formats and has limitations on video length. To address this, we've developed a 3D Slicer Plugin and Gradio API for efficient 3D medical image and video segmentation. We invite you to try them out and provide feedback! 🔧 Deployment: - 3D Slicer Plugin: - Gradio API: (Note: Due to GPU limitations, the online API is available for only 12 hours and may be slow. We highly recommend deploying the Gradio API with your own computing resources: A big shoutout to Jun Ma (JunMa) who recently joined our UHN AI hub (UHN AI Hub) as Machine Learning Lead, and kudos to all co-authors: Sumin Kim, Feifei Li, Mohammed Baharoon (Mohammed Baharoon), Reza Asakereh, and Hongwei Lyu! This is true teamwork! Looking forward to collaborating with the community to advance 3D medical image and video segmentation foundation models! University Health Network U of T Department of Computer Science Department of Laboratory Medicine & Pathobiology Temerty Centre for AI in Medicine (T-CAIREM) Vector Institute #MedTech #AIinHealthcare #DeepLearning #MedicalImaging #SAM2 #MedSAM #AIResearchshow more

Bo Wang
178,481 views • 1 year ago