New Update: Minecraft 3D generation now in under 10s... in falcraft⚡ Text → Image (FLUX.2 Klein) → 3D Gaussian splat (TripoSplat) → Voxelize → Minecraft structure! ~260k colored Gaussians are returned, we filter the faint ones, drop each into a voxel cell, average the color per cell and snap it into the nearest MC block in LAB color spaceshow more

Blendi
31,757 views • 2 months ago
Apple just trained a 3D Gaussian head reconstruction model... on 10,000+ subjects. Feed-forward. No test-time optimization. New identity in, reconstructed Gaussian head out. The UV-parameterized Gaussian representation decouples the number of Gaussians from the number and resolution of input images, making it practical to train with many high resolution views. And the heads are not just static either: text-conditioned identity generation, plus blendshape-driven latent animation across identities. We've been building in the 3D Gaussian Splatting space for a while. The gap between "research demo" and "works on real people at scale" is closing fast.show more

KIRI Engine - 3D Scanner App
12,181 views • 3 months ago
The new Krea 3D feature is insanely good. You... can turn an image into a 3D model to move & rotate in space - and it guides real-time scene generation. Watch me move this couch around a living room ⬇️show more

Justine Moore
32,728 views • 1 year ago
Everyone's sleeping on image-to-3D AI models. They can make... your app look incredibly unique, with just a little effort. Here's how. This is my calorie tracker, built in a week with nothing but prompting. Just Claude Code + a couple APIs. The visuals are all AI-generated. I'll be sharing the full workflow + all the crazy technical stuff Claude and I did to make this work, so nobody has to struggle through it like me. Deep dive coming soon! Till then, this is the high-level idea: 1. Get a clean image of the food (or whatever your asset is) - In my app, the user describes foods via text, or attaches images (or both) - If text, an LLM extracts the food description and formats it into a specific prompt I tuned for this design, and we generate an image using Z-Image Turbo through fal - If image, we do the same thing but with FLUX.2 [dev] to edit the user image into our reference design - Originally, both used Google Nano Banana, but switching to open models cut costs and latency a ton 2. Gaussian splatting (2D image → 3D model) - I tried various 2D-to-3D options on fal and ended up with TripoSplat as my preferred balance of speed, cost, latency; this turns an image into a 3D model that looks super high quality (link below) - The app displays the 2D image while our backend generates the 3D splat - We "groom" the splat to reduce size and load time by culling low-opacity/scale points 3. Render efficiently on device Originally, it looked great but ran at 10 FPS. Getting to 120 FPS was a crazy journey. TL;DR: - SwiftUI had to go; it forced us to render each asset in independent MTKViews, which wasn't workable - Instead, we composite every dish into one full-bleed CAMetalLayer using MetalSplatter (link below) - We had to make some optimizations within MetalSplatter's code too, to reduce the overhead of sorting points per render Then I added some finishing touches like the subtle rotation and parallax as they move around. I think it turned out pretty cool :) Overall, this took some effort, but we still got it done in less than a day. Hopefully your agent can follow in the footsteps of mine and do it much faster. Keep an eye out for the bigger writeup, which'll give your agent everything it needs. If you have any questions, drop em below!show more

Anshu
19,931 views • 2 months ago
I used a movement sheet as a reference image... to animate the dance using Seedance 2.0 + GPT image 2.0 GPT Image 2.0 prompt: [STYLE] Monochrome grayscale illustration, 3D-rendered character, clean instructional reference sheet, white background, comic-style cell grid layout, technical diagram aesthetic. [LAYOUT] 4×4 grid layout with a total of 16 panels. Each panel is separated by thin black border lines. Cells are numbered from 1 to 16, with consistent panel sizes. [CHARACTER] image1 (the same character appears consistently in all panels) [PANEL STRUCTURE – per cell] Top-left: bold number badge + English title text Center: full-body character pose illustration Bottom-left: English description text (3–4 lines) Overlay: directional arrows indicating movement [ARROWS / MOTION INDICATORS] Curved arrows, straight arrows, and circular rotation indicators placed around the character to show motion flow and direction. [RENDERING STYLE] Highly detailed 3D sculpted style, soft studio lighting, subtle shadows, no color, grayscale shading, clean linework, game concept art quality. [NEGATIVE] No background scenery, no color tones, no additional characters, no complex background.show more

Oogie
353,047 views • 4 months ago
GPT Image 2 + Seedance 2.0 I used a... movement sheet as a reference image to animate the dance. Prompt: Monochromatic grayscale illustration, 3D-rendered character, clean instructional reference sheet. White background, comic-style cell grid layout, technical diagram aesthetic. [LAYOUT] 4×4 grid layout (16 panels total). Each panel separated by thin black border lines. Panels are evenly sized and consistently aligned. Each cell is clearly numbered from 1 to 16. [CHARACTER] (Insert character description) Example: Young female dancer with an athletic build, ponytail hairstyle, wearing a crop top, baggy pants, and sneakers. The same character must appear consistently in all panels. [PANEL STRUCTURE – per cell] Top-left: bold number badge + Korean title text Center: full-body character pose illustration Bottom-left: Korean description text (3–4 lines) Overlay: directional arrows indicating movement flow [ARROWS / MOTION INDICATORS] Curved arrows, straight arrows, and circular rotation indicators. Arrows should be placed around the character to clearly show movement direction and flow. [RENDERING STYLE] Highly detailed 3D sculpted style. Soft studio lighting with subtle shadows. No color — grayscale only. Clean linework, polished finish, game concept art quality. [NEGATIVE PROMPT] No background scenery. No color tones. No additional characters. No cluttered or complex backgrounds.show more

K
13,813 views • 4 months ago
How to generate 3D miniature city models and animate... them using Kling AI? This visual effect can be created with Kling O1, and then rotated in 3D using Image to Video. The image prompt used for generation is as follows: Present a clear, 45° top-down isometric miniature 3D cartoon scene of New York featuring its most iconic landmarks and architectural elements. Use soft, refined textures with realistic PBR materials and gentle, lifelike lighting and shadows. Integrate the current weather conditions directly into the city environment to create an immersive atmospheric mood.Use a clean, minimalistic composition with a soft, solid-colored background. At the top-center, place the title "New York" in large bold text in white. You can create different effects by changing the city name according to your needs.show more

Kling AI
34,303 views • 8 months ago
I used a movement sheet as a reference image... to animate the dance using Seedance 2.0 + GPT image 2.0 GPT Image 2.0 prompt: [STYLE] monochromatic grayscale illustration, 3D rendered character, clean instructional reference sheet, white background, comic-style cell grid layout, technical diagram aesthetic [LAYOUT] 4x4 grid layout, 16 panels total, each panel separated by thin black border lines, numbered cells from 1 to 16, consistent panel size [CHARACTER] (캐릭터 설명 입력) 예: young female dancer, athletic build, ponytail hairstyle, crop top and baggy pants, sneakers, same character in all panels [PANEL STRUCTURE - per cell] top-left: bold number badge + Korean title text center: full-body character pose illustration bottom-left: Korean description text (3-4 lines) overlay: directional arrows indicating movement direction [ARROWS / MOTION INDICATORS] curved arrows, straight arrows, circular rotation indicators, placed around the character to show movement flow and direction [RENDERING STYLE] high detail 3D sculpt style, soft studio lighting, subtle shadows, no color, grayscale shading, clean linework, game concept art quality [NEGATIVE] no background scenery, no color tones, no extra characters, no cluttered backgroundsshow more

Ciri
146,410 views • 4 months ago
Claude Opus 4.8 + OpenClaw now turns a restaurant's... menu photos into 3D models guests can view, customize and order, then mails the owner a postcard with the QR…on autopilot. here's how agencies can land recurring contracts with this system: - scans every restaurant in a city in real time - pulls their real reviews, ratings, and reviewer-uploaded food photos - turns each dish photo into an interactive 3D model - samples the brand color straight from the restaurant's photos - guests can view each dish, customize it, and order on the spot - writes a printed postcard about their best dish - mails it to the registered office, addressed to the owner, with a QR to the live 3D menu every step from the scrape to the 3D models to the mailbox is automated reply "3D" + RT and i'll send you a free guide so you can build this too (must be following so i can DM you)show more

Chris
11,853 views • 2 months ago
A sneak peak of a complex and technically challenging... experiment that my lab developed: Super proud of PhD candidate Hannah Johnson for showcasing our Whole-gut spatial genomic analysis in #zebrafish. This video illustrates one landmark in the protocol after multiple rounds of sequential #HCR and 3D imaging in zebrafish larvae to reveal spatial expression of numerous mRNAs in the same specimen. Data from these imaging data sets are then computationally analyzed for spatial cell groups, spatially variable genes, and differentially expressed genes along 3D. We are leveraging this systems-level SGA to uncover unappreciated mechanistic insight at the cell and tissue levels into #ENS construction. Stay tuned for our work that exploits this pipeline within various mutant and perturbation conditions. Reach out if you are interested in trying this! #fruitypebblesshow more

Rosa Uribe, PhD
10,119 views • 6 months ago
HTML enters 3D! Or vice versa? With the new... HTML in Canvas by WICG, we can finally put native DOM elements directly into WebGL/WebGPU scenes. It is experimental for now, but the possibilities for 3D interfaces and special effects are huge. This demo was built using Three.js and Omma AI (tool by Spline ) It’s a fun new way to explore what the web can do! Are you interested in seeing the demo?show more

Gábor Pribék
176,349 views • 4 months ago
🎨 Gaussian Splat rigging opens infinite artistic possibilities! Realtime... rendering with my VHS effect in VR 1.5M splats running at 80 FPS in VR with bone-based Mixamo rig animation. Travel into the volumetric worlds of Valerian and the City of a Thousand Planets, reimagined in that bold 1980s sci-fi aesthetic ✨ World Labs Theoretically Media Technical • GPU compute shader skinning (256 threads) • Linear Blend Skinning (LBS) with up to 4 bones per splat • Bind pose capture for accurate deformation • CommandBuffer synchronization for frame-perfect timing Why splats are perfect for comic art I think: Each splat is a volumetric 3D ellipsoid that naturally creates crisp edges and flat color regions—the essence of cel-shading. No complex shader tricks needed. The volumetric representation itself gives you that graphic novel aesthetic. #Realtime #GaussianSplatting #VR #GameDev #Unity3D #VolumetricRendering #ComicArtshow more

Daniel Skaale
15,155 views • 9 months ago
SOMEONE FROM TOKYO IS MAPPING BIRD LANGUAGE INTO REAL... DATA PATTERNS AND THE VISUALIZATION LOOKS LIKE A NEURAL NETWORK DREAMING Every bird sound - frequency, duration, amplitude and modulation - gets converted into 3D space coordinates and rendered as a cluster of colored points in real time through Deepen AI. A microphone captures the acoustic signal, FFT breaks it down into frequency components from 1 kHz to 8 kHz where most birds communicate, an ML model classifies the pattern and the mapped data hits a graph where each call type gets its own color and position in space. Birds have a vocal repertoire of 5 to 200+ unique signals depending on the species - and every signal carries different information about predators, food, territory and mates that humans simply can't hear. The same technology that decodes bird language detects anomalies in industrial machinery, cardiac rhythms and structural vibrations in buildings - he just trained it on nature first.show more

Cortex
11,128 views • 3 months ago
The most detailed 3D reconstruction of a cell ever... created. Blows my mind every time. But what exactly are we looking at here? The average human cell contains: ~ 15-20 total distinct organelle types, totalling between ~1-10 million working together per cell. All these nano-machines in the cell are made up of proteins. ~ 8,000-10,000 distinct types of unique proteins, adding up to between 40 million - 10 trillion total proteins making up all those cellular systems. ~ 10,000 - 15,000 distinct types of RNA shuttling information around the cell, totalling up to ~10 million RNA molecules moving around the cell simultaneously. ~ Billions of Lipid molecules packed together into the cell membrane, which is also packed tightly with millions more protein-based nano-machines. And let's not forget billions of lines of DNA information to build and run it all. That's TRILLIONS of of individual molecular pieces working together to make a single cell function. That means there is more complexity in a single cell than humanity's largest cities. And people still believe this wasn't Divinely Designed. This is God's Glory on Display. But to make the point. A cell couldn't have evolved from some nebulous simpler "protocell" because even the simplest cells still require massive complexity. The "simplest" cell ever created was engineered by scientists knocking out pieces of a functional cell until it stopped functioning. Here is what they found is the absolute necessary minimal requirements of a cell to function: - Over ~531,000 lines of coded DNA information - 473 total genes to create hundreds of unique protein products (they later added 19 genes back in because the cell was so weak) - Hundreds of thousands of total proteins all working together - Extensive regulatory networks guiding all these interactions If the cell doesn't have all these systems in place, from the start... it doesn't live. Cell rely on an intricate network of complex systems, which are themselves built from complex interconnected pieces woven together into an incomprehensibly complex web of functionilty. Only intelligence has ever been observed creation vast interconnected systems like this. Life was clearly Created. It couldn't happen any other way.show more

Divinely Designed
166,325 views • 3 months ago
STEVE-1: A Generative Model for Text-to-Behavior in Minecraft paper... page: Constructing AI models that respond to text instructions is challenging, especially for sequential decision-making tasks. This work introduces an instruction-tuned Video Pretraining (VPT) model for Minecraft called STEVE-1, demonstrating that the unCLIP approach, utilized in DALL-E 2, is also effective for creating instruction-following sequential decision-making agents. STEVE-1 is trained in two steps: adapting the pretrained VPT model to follow commands in MineCLIP's latent space, then training a prior to predict latent codes from text. This allows us to finetune VPT through self-supervised behavioral cloning and hindsight relabeling, bypassing the need for costly human text annotations. By leveraging pretrained models like VPT and MineCLIP and employing best practices from text-conditioned image generation, STEVE-1 costs just $60 to train and can follow a wide range of short-horizon open-ended text and visual instructions in Minecraft. STEVE-1 sets a new bar for open-ended instruction following in Minecraft with low-level controls (mouse and keyboard) and raw pixel inputs, far outperforming previous baselines. We provide experimental evidence highlighting key factors for downstream performance, including pretraining, classifier-free guidance, and data scaling. All resources, including our model weights, training scripts, and evaluation tools are made available for further research.show more

AK
144,806 views • 3 years ago
This is some quietly impressive work on making video... world models actually controllable in 4D space. VerseCrafter lets you take an input image, use something like Blender to animate the 3D camera path and object trajectories, then uses that to condition generation. Scribbling in 2D feels so crude in comparison. The authors represent everything in a shared 4D world state - static background as a point cloud, moving objects as 3D gaussian trajectories. The gaussians are an interesting choice because they capture position, shape, and orientation probabilistically rather than forcing rigid bounding boxes or category specific models like SMPL-X for human bodies. They bolt this onto frozen Wan2.1 with a lightweight adapter, so they get a strong video prior. They also built a pipeline to auto extract 4D annotations from real world videos to train this puppy. It doesn't look sexy yet, but IMO this is the interface video world models need - actual 3D authoring tools to exert control rather than crude scribbles and prompt incantations.show more

Bilawal Sidhu
26,017 views • 7 months ago
✨ I can now generate 3d assets for my... drone sim at directly from Cursor (sponsor of #vibejam) I need buildings that you'd see in a war torn city, like warehouses in ruins, broken down abandoned houses, bombed out bridges etc. Nano Banana Pro or 2 can generate them really well and then you can put them in an image-to-3d model and you get a GLB or FBX That one you can then import into your Three.js game, the models might be big though, in my case like 16MB, so I ask it to compress it and make it more low poly so it loads fast ThreeJS then loads the individual GLBs on page load and puts them in my drone sim somewhere randomly, I think I should remove some of the grass and match the sandy color of the ruins though to make it fit in moreshow more

@levelsio
134,777 views • 4 months ago
Is Google taking initial steps to enhance Street View?... For some reason, Street View seems stuck in technology that feels outdated. I wonder if we'll see such improvements on the product side. Also, note how much better it performs in all aspects compared to Zip-NeRF in their presented material. It offers more details and fewer artifacts. Great work! "LODGE: Level-of-Detail Large-Scale Gaussian Splatting with Efficient Rendering" Contributions: • We propose a novel LOD representation for 3DGS which, unlike previous methods [27, 28, 17], does not recompute the list of used Gaussians at each frame. This allows for acceleration and compaction, enabling the rendering of large-scale scenes even on mobile devices. • We design a strategy to automatically select optimal hyperparameters for splitting LODs, whereas most other methods require manual tuning of hyperparameters for each 3D scene. • To further accelerate rendering, we split the scene into chunks and pre-compute sets of active Gaussians per chunk. • Finally, we introduce a novel opacity interpolation scheme to produce visually pleasing rendering and eliminate artifacts when transitioning between chunks.show more

MrNeRF
62,564 views • 1 year ago
With Hunyuan3D World Model 1.0 now released and open-sourced,... we're excited to showcase the technical highlights behind this impressive innovation: ✅360° Panoramic Generation: Creates complete, immersive “world scenes”, far beyond localized views. ✅Explorable 3D Scene Generation: Generates diverse, spatially consistent 3D worlds from text/image for truly immersive exploration. ✅Interactive/Editable: Achieves separation of foreground objects, background terrain, ground, and sky, for seamless secondary editing. ✅Exportable Mesh: Generated scenes can be exported as 3D meshes for direct import into mainstream game engines and modeling software. ✅Industry-Leading SOTA Evaluation: Surpasses state-of-the-art open-source models in generation quality. As the industry's first open-source model for physical simulation and explorable world generation, Hunyuan3D World Model 1.0 aims to foster a collaborative community ecosystem with developers and enthusiasts. ✨ Try it now: 🤗 Hugging Face:show more

Tencent Hy
23,203 views • 1 year ago
🚨PHYSICS NEWS🚨: Scientists Finally Complete Schrödinger’s 100-Year-Old Color Theory... — Uniphics Reveals the Deeper Structure Behind Perception 🧨 On June 7, 2026, researchers at Los Alamos National Laboratory announced that they have finally resolved a long-standing problem in Erwin Schrödinger’s 1920s theory of color perception. Their work shows that the way humans experience qualities like hue, saturation, and brightness is not arbitrary but emerges directly from the mathematical structure of color space itself. **Uniphics provides a natural explanation for why color perception has such a precise underlying geometry.** In Uniphics, what we experience as color ultimately arises from how light (electron spin waves) interacts with matter within the ξM-field. When light encounters different materials, the spin-wave patterns are modified in specific ways depending on the local energy density and the arrangement of Gyrotrons in the material. These modifications create distinct interference patterns that our visual system then interprets as different colors. The fact that color perception follows clean mathematical rules — as Schrödinger suspected and Los Alamos researchers have now confirmed — makes sense in Uniphics because the underlying spin interactions and energy density gradients are themselves highly structured. Negentropy favors organized, low-energy configurations, which leads to consistent and predictable ways that spin waves are altered by different materials. This produces the orderly geometry of perceived color space rather than random or chaotic sensations. In this view, the mathematics of color is not just a useful model — it reflects real physical organization in the ξM-field. The qualities we perceive as color are downstream effects of how spin waves propagate and interfere under varying energy density conditions. This breakthrough in understanding color perception is another example of how fundamental organizing principles can explain phenomena that once seemed mysterious or purely subjective. Could many other aspects of human perception ultimately trace back to the same energy density and spin-wave dynamics that govern the physical world? **A Theory of Everything should be able to answer everything.** Uniphics Explained Simply PDF: Chapters 1–10 free: Grokipedia: #Uniphics #TheoryOfEverything #ColorPerception #Physics #LosAlamos Grok xAIshow more

Paul Maley
38,808 views • 2 months ago