Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Turn a single 2D image into a fully interactive, animated 3D character - running directly in your browser.🥊 This is the Ringside Boxer, built with img2threejs v1.5.1. Live showcase: 2D Image → 3D Reconstruction (TripoAI/Hyper3D) → Reference GLB + Specs → Rig & Animation Setup → Animation Loops →...

45,937 görüntüleme • 4 gün önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

Want to create an avatar from a single image? FlexAvatar is a transformer model that creates full 360°, high-quality, and expressive 3D head avatar from just a single portrait image in minutes. Real-time Demo: FlexAvatar's lightweight architecture allows both animation and rendering in real-time, enabling interactive user experiences. To create a new 3D head avatar, only one image is required, e.g., from a webcam. The final avatar is ready after 2 minutes. Architecture: Under the hood, FlexAvatar adopts a transformer-based encoder-decoder design. The encoder maps the input image onto a latent avatar space, while the decoder produces 3D Gaussian attribute maps by incorporating the animation signal via cross-attention. The model learns all facial animations directly from the data without relying on pre-built 3D face models. This equips the avatars with realistic facial expressions. The internal avatar latent space can be conveniently used to integrate additional observations of a person via fitting. This enables use-cases where more than one image of a person is available, e.g., from a phone scan of the person. We train jointly on 2D monocular videos and multi-view data. However, in monocular videos, the animation signal leaks the target viewpoint, causing the model to produce incomplete 3D heads. We call this phenomenon entanglement of driving signal and target viewpoint. To prevent entanglement, we introduce bias sinks. These are learnable tokens that indicate whether a training sample stems from a monocular or a multi-view dataset. During training, the model learns to produce incomplete 3D heads only when the monocular token is present. During inference, FlexAvatar then always uses the multi-view token for which the model has learned to produce complete 3D heads. This simple design allows to combine the generalizability from monocular data with the quality of multi-view data. FlexAvatar summary: - Input: Single-image, phone scan, or monocular video - Output: Full 360° head avatar - Expressive animations - Real-time rendering and animation - Generalization to any portrait - Create a new avatar in 2 minutes - Use bias sinks to combine 2D and 3D data 🏠 🌍 🎥 Great work by Tobias Kirschstein and Simon Giebenhain!

Matthias Niessner

96,238 görüntüleme • 8 ay önce

We just shipped a completely new concept for web interaction. Live on three․ws, we are thrilled to showcase this demo of persistent, interactive 3D AI Agents that live right in your browser. As the first mover in this space, we are achieving what hasn't been done before: seamlessly bridging the gap between traditional flat web pages and fully immersive, spatial 3D environments. We are fundamentally shifting how humans experience the internet. Here's how. Intelligent AI Agents. Your 3D companion drops right into the screen, tracks your cursor, and interacts with you based on the exact page you are viewing. Context-Aware Guide. Knows exactly where visitors are. It can point out and encourage clicks on the most important parts of a site. Highly Interactive. Turns to follow your cursor in real-time, waves when you navigate, and idles naturally. Playground Mode. Click your avatar and it seamlessly detaches into a full-page 3D stroll or platformer right in the browser. Use keyboard arrows on desktop and joystick on mobile. Choose Your Avatar. Pick exactly who walks the web with you from a fully customizable, hot-swappable roster. Polite & Lightweight. Built with smart compression and shared animation files. It monitors browser memory to guarantee zero slowdowns and respects "reduced motion" accessibility settings. Flat websites will be gone before you know it. The future is interactive, the future is intelligent. It's time to break AI out the chatbox. three․ws

three.ws

22,706 görüntüleme • 2 ay önce

🚀 Announcing Echo — our new frontier model for 3D world generation. Echo turns a simple text prompt or image into a fully explorable, 3D-consistent world. Instead of disconnected views, the result is a single, coherent spatial representation you can move through freely. This is part of a bigger shift in AI: from generating pixels and tokens to generating spaces. Echo predicts a geometry-grounded 3D scene at metric scale, meaning every novel view, depth map, and interaction comes from the same underlying world — not independent hallucinations. Once generated, the world is interactive in real time. You control the camera, explore from any angle, and render instantly — even on low-end hardware, directly in the browser. High-quality 3D world exploration is no longer gated by expensive equipment. Under the hood, Echo infers a physically grounded 3D representation and converts it into a renderable format. For our web demo, we use 3D Gaussian Splatting (3DGS) for fast, GPU-friendly rendering — but the representation itself is flexible and can be easily adapted. Why this matters: consistent 3D worlds unlock real workflows — digital twins, 3D design, game environments, robotics simulation, and more. From a single photo or a line of text, Echo builds worlds that are reliable, editable, and spatially faithful. Echo also enables scene editing and restyling. Change materials, remove or add objects, explore design variations — all while preserving global 3D consistency. Editing no longer breaks the world. This is only the beginning. Echo is the foundation for future world models with dynamics, physical reasoning, and richer interaction — environments that don’t just look right, but behave right. Explore the generated worlds on our website and sign up for the closed beta. The era of spatial intelligence starts here. 🌍 #Echo #WorldModels #SpatialAI #3DFoundationModels Check it out:

SpAItial AI

176,524 görüntüleme • 8 ay önce

Only if education could be this interactive ❤️‍🔥 I've had a looong wish to build something genuinely useful through vibe coding, and I finally did it. A 3D human anatomy application built with Three.js using GPT 5.6 Sol. It all started with a single design image that I created using GPT Image 2.0. I then used it to generate every 3D organ image, one by one. Next, I converted each of those images into 3D models using Tripo (and no, they didn't sponsor this 😄). After that, I opened Codex, wrote a master prompt based on the design, and gave it the prompt, the design image, and all the 3D models. Codex built the first version beautifully, but there was one big problem. Every single 3D model was nearly 120-150 MB. That obviously wasn't practical for the web and was giving a performance of 16fps. After a few iterations, Codex optimized each model down to roughly 2–5.5 MB while preserving the visual quality, reducing the total asset size from ~900 MB to just 28.6 MB. And each model loads on demand. Along the way, Codex also generated those anatomical illustrations showing where each organ sits in the human body, and even created the interactive hotspot markers that explain different parts of every organ. It handled all of that. The process wasn't exactly one shot, but it also wasn't difficult. You just have to do it step by step. It genuinely felt like building something that could make learning anatomy much more engaging. The inspiration came from Dilum Sanjaya's 3D animal plant cell project. I remember seeing it and thinking, "I want to build something like this one day." And I did it :D Live: Code:

The Bugged Dev

2,088,873 görüntüleme • 1 ay önce

So Runway Gen 4.5 finally adds image-to-video, the workflow most pros rely on for consistency. We put it head to head with Kling AI and Flow by Google VEO using the same reference images and prompts (below) to evaluate motion quality, stability, and cinematic realism. 1. Action/WaterPhysics Test Prompt: Cinematic, wide-shot of a man running in a shallow river. The camera is tracking the man from behind as he runs up the river. Handheld camera shake as the camera follows the man. 2. Fire Physics Test Prompt: Cinematic, wide-shot of terrified woman running towards her burning barn. She abruptly stops, and puts in hands on her head as she watches her barn burn down. 3. VFX test prompt: Cinematic, wide-shot of a hooded figure. Flashes of purple magic and smoke whirl around the figure. The figure lifts its arms as the purple magic and smoke intensifies. 4. 2D Animation Test Prompt: 2D animated shot of a waiting at a bus stop in a thunderstorm. The man turns, walks to the bench, and sits down. 5. 3D Animation Test Prompt: 3D animated shot of an octopus. The octopus reaches into a coral and picks up a glowing white gem. 6. Conversation Test Prompt: slow camera push-in as two friends are having a conversation at a coffee shop Overall verdict: Despite the “world’s best” claim, Runway Gen 4.5 is not there yet. Prompt adherence is solid, but motion, physics, and cinematic realism still lag behind tools like Kling and VEO. Great platform, mid-tier model for now.

Curious Refuge

25,538 görüntüleme • 7 ay önce

When I saw the mask "Tribes of the Calf" from Kanbas I knew I had to make it into reality. The jewelry and gold really made it stand out for me. Since Sam Spratt's The Masquerade was revealed, I have been spending time sculpting and dissecting the mask to recreate it in 3D as faithfully as possible. I delved into the creation of this mask for many reasons. I love a good challenge and this mask surely was one for me. Creating something in 3D from a 2D image is not easy, and especially when the source has generative nature, some stuff is hard to interpret, but I tried my best to make sure the visual integrity of the mask is as close to the original as possible. Splitting the whole mask into parts, filling the missing pieces so I can build the textures was quite a lot of work. I tried to present the mask in my own style with a slightly different colorway to adapt to the mask itself. Please enjoy this short animation, and turn on sound🔊 This piece is my statement that I am here to stay. That I have a voice that often feels being lost in the void. That I have been creating and posting digital art for over 20 years now and will continue until I'm gone. I have a story to tell and I want to be heard. The space we have here is small, and is shrinking day by day. It doesn't have to be like that. We need to support each other and push ourselves and people here, otherwise we are all doomed. As Kanbas has put in their observation of the mask: "Inspirational. Emotional. Natural." This is what our space can be, and this is me making a statement with this homage. I will share a 4k still below as well as a short video showing the 3D GLB interactive model together with a yt link to the 4k video since compression here is pretty bad.

shoneec

17,894 görüntüleme • 1 yıl önce