Loading video...

Video Failed to Load

Go Home

New Update: Minecraft 3D generation now in under 10s in falcraft⚡ Text → Image (FLUX.2 Klein) → 3D Gaussian splat (TripoSplat) → Voxelize → Minecraft structure! ~260k colored Gaussians are returned, we filter the faint ones, drop each into a voxel cell, average the color per cell and snap...

31,917 views • 3 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Everyone's sleeping on image-to-3D AI models. They can make your app look incredibly unique, with just a little effort. Here's how. This is my calorie tracker, built in a week with nothing but prompting. Just Claude Code + a couple APIs. The visuals are all AI-generated. I'll be sharing the full workflow + all the crazy technical stuff Claude and I did to make this work, so nobody has to struggle through it like me. Deep dive coming soon! Till then, this is the high-level idea: 1. Get a clean image of the food (or whatever your asset is) - In my app, the user describes foods via text, or attaches images (or both) - If text, an LLM extracts the food description and formats it into a specific prompt I tuned for this design, and we generate an image using Z-Image Turbo through fal - If image, we do the same thing but with FLUX.2 [dev] to edit the user image into our reference design - Originally, both used Google Nano Banana, but switching to open models cut costs and latency a ton 2. Gaussian splatting (2D image → 3D model) - I tried various 2D-to-3D options on fal and ended up with TripoSplat as my preferred balance of speed, cost, latency; this turns an image into a 3D model that looks super high quality (link below) - The app displays the 2D image while our backend generates the 3D splat - We "groom" the splat to reduce size and load time by culling low-opacity/scale points 3. Render efficiently on device Originally, it looked great but ran at 10 FPS. Getting to 120 FPS was a crazy journey. TL;DR: - SwiftUI had to go; it forced us to render each asset in independent MTKViews, which wasn't workable - Instead, we composite every dish into one full-bleed CAMetalLayer using MetalSplatter (link below) - We had to make some optimizations within MetalSplatter's code too, to reduce the overhead of sorting points per render Then I added some finishing touches like the subtle rotation and parallax as they move around. I think it turned out pretty cool :) Overall, this took some effort, but we still got it done in less than a day. Hopefully your agent can follow in the footsteps of mine and do it much faster. Keep an eye out for the bigger writeup, which'll give your agent everything it needs. If you have any questions, drop em below!

Anshu

29,342 views • 3 months ago

The most detailed 3D reconstruction of a cell ever created. Blows my mind every time. But what exactly are we looking at here? The average human cell contains: ~ 15-20 total distinct organelle types, totalling between ~1-10 million working together per cell. All these nano-machines in the cell are made up of proteins. ~ 8,000-10,000 distinct types of unique proteins, adding up to between 40 million - 10 trillion total proteins making up all those cellular systems. ~ 10,000 - 15,000 distinct types of RNA shuttling information around the cell, totalling up to ~10 million RNA molecules moving around the cell simultaneously. ~ Billions of Lipid molecules packed together into the cell membrane, which is also packed tightly with millions more protein-based nano-machines. And let's not forget billions of lines of DNA information to build and run it all. That's TRILLIONS of of individual molecular pieces working together to make a single cell function. That means there is more complexity in a single cell than humanity's largest cities. And people still believe this wasn't Divinely Designed. This is God's Glory on Display. But to make the point. A cell couldn't have evolved from some nebulous simpler "protocell" because even the simplest cells still require massive complexity. The "simplest" cell ever created was engineered by scientists knocking out pieces of a functional cell until it stopped functioning. Here is what they found is the absolute necessary minimal requirements of a cell to function: - Over ~531,000 lines of coded DNA information - 473 total genes to create hundreds of unique protein products (they later added 19 genes back in because the cell was so weak) - Hundreds of thousands of total proteins all working together - Extensive regulatory networks guiding all these interactions If the cell doesn't have all these systems in place, from the start... it doesn't live. Cell rely on an intricate network of complex systems, which are themselves built from complex interconnected pieces woven together into an incomprehensibly complex web of functionilty. Only intelligence has ever been observed creation vast interconnected systems like this. Life was clearly Created. It couldn't happen any other way.

Divinely Designed

168,535 views • 4 months ago

STEVE-1: A Generative Model for Text-to-Behavior in Minecraft paper page: Constructing AI models that respond to text instructions is challenging, especially for sequential decision-making tasks. This work introduces an instruction-tuned Video Pretraining (VPT) model for Minecraft called STEVE-1, demonstrating that the unCLIP approach, utilized in DALL-E 2, is also effective for creating instruction-following sequential decision-making agents. STEVE-1 is trained in two steps: adapting the pretrained VPT model to follow commands in MineCLIP's latent space, then training a prior to predict latent codes from text. This allows us to finetune VPT through self-supervised behavioral cloning and hindsight relabeling, bypassing the need for costly human text annotations. By leveraging pretrained models like VPT and MineCLIP and employing best practices from text-conditioned image generation, STEVE-1 costs just $60 to train and can follow a wide range of short-horizon open-ended text and visual instructions in Minecraft. STEVE-1 sets a new bar for open-ended instruction following in Minecraft with low-level controls (mouse and keyboard) and raw pixel inputs, far outperforming previous baselines. We provide experimental evidence highlighting key factors for downstream performance, including pretraining, classifier-free guidance, and data scaling. All resources, including our model weights, training scripts, and evaluation tools are made available for further research.

AK

144,811 views • 3 years ago