Loading video...

Video Failed to Load

Go Home

All these consistent 3D characters were made with AI in under 20 min via a fast, simple workflow 🧵👇 🛠️ One tool: - Custom LoRA - Flux Kontext (back view control) - 3D gen: Rodin / Tripo / Hunyuan ⏱️ Get clean meshes + PBR textures in record time.

67,916 views • 1 year ago •via X (Twitter)

10 Comments

Emm | scenario.com's profile picture
Emm | scenario.com1 year ago

Step 1 - "Front View" Generation Start by generating the front view of each character (with a custom LoRA to lock in the style) I used “Action Plastic", a model designed to create characters with the toy-like aesthetic of classic action figures Link 👉

Emm | scenario.com's profile picture
Emm | scenario.com1 year ago

Step 2 - "Back View" Once the front view is ready, generate the back. Head over to and select Flux Kontext ✏️ Prompt: “Generate the back view of this character” (feel free to add details) This step gives you control over front/back before going 3D.

Emm | scenario.com's profile picture
Emm | scenario.com1 year ago

Step 3 - Click “Convert to 3D” on any character you like. Choose your preferred AI model from 7 different generators (Rodin, Tripo, Hunyuan, etc.). For better accuracy, load the back view from in Step 2. Then just hit Generate, and get an accurate textured 3D mesh in minutes.

Emm | scenario.com's profile picture
Emm | scenario.com1 year ago

Step 4 - Compare Models Instantly If not happy with the first result: 🔁 Switch the 3D model in one click - no need to re-upload anything or switch app. 🆚 Generate outputs side by side to compare mesh and texture fidelity. Quick, flexible iteration to find the best fit.

Emm | scenario.com's profile picture
Emm | scenario.com1 year ago

Step 5 – Export the full collection Once your your character set is ready: 📥 Batch select and download all your 3D assets at once into Blender or any 3D software of your choice for further editing. From idea to a full pack of editable 3D assets... in under 20 minutes.

Emm | scenario.com's profile picture
Emm | scenario.com1 year ago

+ Enjoy automatic indexing of all assets, smart search and filters (type, author, date, model, collection, tags…) - so you can find past work in seconds, even months later. SOC2 compliant. SSO/SAML & enterprise features available. More details 👇

Ben Pielstick's profile picture
Ben Pielstick1 year ago

Still looks like uniform topology though. Probably good for 3D printing and not much else?

Emm | scenario.com's profile picture
Emm | scenario.com1 year ago

AI meshes are dense, it's no secret 🫣. You can tweak mesh density in the model settings - down to 500 faces or less with Hunyuan 2.1 on simple objects, for ex. For more complex assets, bring them into Blender for cleanup/retopo (until AI retopology makes it faster & easier)

xenoshiba - robotmasters.io 🎮's profile picture
xenoshiba - robotmasters.io 🎮1 year ago

wishing there will be a proper workflow for lowpoly models soon 🤞

Emm | scenario.com's profile picture
Emm | scenario.com1 year ago

How about these?

Related Videos

🚀 Announcing Echo — our new frontier model for 3D world generation. Echo turns a simple text prompt or image into a fully explorable, 3D-consistent world. Instead of disconnected views, the result is a single, coherent spatial representation you can move through freely. This is part of a bigger shift in AI: from generating pixels and tokens to generating spaces. Echo predicts a geometry-grounded 3D scene at metric scale, meaning every novel view, depth map, and interaction comes from the same underlying world — not independent hallucinations. Once generated, the world is interactive in real time. You control the camera, explore from any angle, and render instantly — even on low-end hardware, directly in the browser. High-quality 3D world exploration is no longer gated by expensive equipment. Under the hood, Echo infers a physically grounded 3D representation and converts it into a renderable format. For our web demo, we use 3D Gaussian Splatting (3DGS) for fast, GPU-friendly rendering — but the representation itself is flexible and can be easily adapted. Why this matters: consistent 3D worlds unlock real workflows — digital twins, 3D design, game environments, robotics simulation, and more. From a single photo or a line of text, Echo builds worlds that are reliable, editable, and spatially faithful. Echo also enables scene editing and restyling. Change materials, remove or add objects, explore design variations — all while preserving global 3D consistency. Editing no longer breaks the world. This is only the beginning. Echo is the foundation for future world models with dynamics, physical reasoning, and richer interaction — environments that don’t just look right, but behave right. Explore the generated worlds on our website and sign up for the closed beta. The era of spatial intelligence starts here. 🌍 #Echo #WorldModels #SpatialAI #3DFoundationModels Check it out:

SpAItial AI

176,105 views • 7 months ago

Alibaba presents MIMO Controllable Character Video Synthesis with Spatial Decomposed Modeling Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case training, which severely limits their applicability of modeling arbitrary characters in a short time. Recent 2D methods break this limitation via pre-trained diffusion models, but they struggle for pose generality and scene interaction. To this end, we propose MIMO, a novel framework which can not only synthesize character videos with controllable attributes (i.e., character, motion and scene) provided by simple user inputs, but also simultaneously achieve advanced scalability to arbitrary characters, generality to novel 3D motions, and applicability to interactive real-world scenes in a unified framework. The core idea is to encode the 2D video to compact spatial codes, considering the inherent 3D nature of video occurrence. Concretely, we lift the 2D frame pixels into 3D using monocular depth estimators, and decompose the video clip to three spatial components (i.e., main human, underlying scene, and floating occlusion) in hierarchical layers based on the 3D depth. These components are further encoded to canonical identity code, structured motion code and full scene code, which are utilized as control signals of synthesis process. The design of spatial decomposed modeling enables flexible user control, complex motion expression, as well as 3D-aware synthesis for scene interactions. Experimental results demonstrate effectiveness and robustness of the proposed method.

AK

148,998 views • 1 year ago