Loading video...

Video Failed to Load

Go Home

🔥Spatial intelligence needs fast, *interactive* 3D world generation 🎮 — introducing WonderWorld: generating 3D scenes interactively following your movement and content requests, and see them in <10 seconds! 🧵1/6 Web: arXiv:

169,131 views • 1 year ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

📢 Our lab has been exploring 3D world models for years — and we’re thrilled to share **PhysTwin**: a milestone that reconstructs object appearance, geometry, and dynamics from just a few seconds of interaction! Led by the amazing Hanxiao Jiang 👉 PhysTwin combines **Gaussian splatting** with **inverse dynamics optimization** based on simple **spring-mass** systems. ⚙️ The result? Real-time, action-conditioned 3D video prediction under novel interactions (i.e., 3D world models). 🔑 A few key takeaways: 1. Having the right structure (e.g., particles/masses) helps navigate the trade-off between sample efficiency, generalization, and broad applicability. 2. Visual foundation models (VFMs) have matured to the point where they can provide rich supervision for world modeling (e.g., tracking, shape completion). 3. Beyond VFMs, many crucial components have come together in recent years: Gaussian splats for rendering, NVIDIA Warp for high-performance simulation, and scene/asset generation from a wide range of labs and companies. The future of 3D world models is looking bright! ✨ 4. The resulting digital twin supports a wide range of downstream applications—especially in data generation and policy evaluation, thanks to its realistic rendering and simulation capabilities. 🎥 All code and data to reproduce the results, along with interactive demos, are available on the website. Check the following visualizations of: (1) observations, (2) reconstructed state/actions, (3) interactive digital twins, and (4) the overlays between real-world robot teleoperation and our model’s open-loop predictions.

Yunzhu Li

25,279 views • 1 year ago

World Model is trending— let's revisit our HunyuanWorld journey. We’ve been pioneering open-source 3D world generation in the past two months, and this ride’s only getting started. 🌍 📅 July: HunyuanWorld 1.0 📌 First open-source 3D world model compatible with CG pipelines (Unity/Unreal/Blender) 📌 Hit 2K+ GitHub stars in just two months ⭐—thank you for the love! 📅 August: 1.0-Lite 📌Same top-tier quality, running on consumer GPUs! 📅 September: 1.0-Voyager 📌 Direct 3D output + world memory—taking exploration further! Seamlessly integrated into CG pipelines with layered 3D modeling (assets, terrain, skybox) and fully open-sourced.. we’re fully committed to building open-source spatial intelligence for all! 🚀 💡 Why it matters? ✅ Seamless CG Pipeline Integration: Export generated 3D scenes as standard mesh formats, effortlessly integrating into industry-standard tools like Blender, Unity, and Unreal Engine for direct editing, animation, and physical simulation. ✅ Hierarchical Scene Editing: Deconstruct scenes into semantic layers (sky, background, foreground objects) via instance recognition and layer decomposition, allowing for atomic-level control—independently modify, relocate, or replace objects without rebuilding the entire world. Project page: Github: Amazing creations by Stijn Spanhove camenduru GENEL | AIを用いた動画制作 apolinario 🌐 とりにく Directive Creator 🪥 👇 #AI #3DGeneration #OpenSource #WorldModels #Hunyuan3D #HunyuanWorld

Tencent HY

20,178 views • 10 months ago

Introducing Kaleido💮 from AI at Meta — a universal generative neural rendering engine for photorealistic, unified object and scene view synthesis. Kaleido is built on a simple but powerful design philosophy: 3D perception is a form of visual common sense. Following this idea, we formulate rendering purely as a sequence-to-sequence generation problem, successfully unifying neural rendering with the architecture principles behind modern language and video models. Unlike traditional neural rendering methods, Kaleido learns 3D purely in a data-driven way, without explicit 3D representations or structures. It acquires spatial understanding directly through large-scale video pretraining, then multi-view 3D data finetuning, inspired by how LLMs acquire textual common sense from large corpora before specialising in domains like coding. Through extensive ablations, we progressively modernised the architecture design and training strategies and tackled key scaling challenges in sequence-to-sequence generative rendering, arriving at a design that’s simple, versatile, and scalable. Kaleido significantly outperforms prior generative models in few-view settings, and remarkably is the first zero-shot generative method matches InstantNGP-level rendering quality in multi-view settings. We view Kaleido also as an alternative step towards world modeling that flexibly spans a spectrum of “realities": with many views, it faithfully reconstructs grounded reality; with fewer views, it imagines plausible unseen details. 🔗 Explore more results and paper:

Shikun Liu

22,332 views • 9 months ago

We ranked a B2B SaaS brand #1 on ChatGPT for their category in 7 days. (And it's being used by marketing teams at Webflow, Chime, and Deepgram) This platform tracks AI visibility + generates cited content automatically across ChatGPT, Perplexity, Claude, and Gemini... → No more 6-12 months waiting for Google rankings to move → No more $60K agency dashboards that only show problems → No more 10 different tools to track, create, and publish content → No more manual content gap analysis taking 20+ hours weekly → No more AI slop that ChatGPT refuses to cite Just connect your data sources → autonomous visibility tracking + content generation system. Here's how it works: → AI Citation Scanner (tracks mentions across ChatGPT, Perplexity, Claude, Gemini) → Competitive Gap Analysis (identifies where competitors get cited and you don't) → First-Party Data Integration (connects Zendesk, HubSpot, Drive, product docs) → AI Content Generator (creates authoritative content with human review checkpoints) → Direct CMS Publishing (publishes to Webflow, Contentful automatically) → Performance Measurement (tracks results across traditional + AI search) Companies using this infrastructure: • Webflow: 40% traffic lift + 5X content velocity • Chime: 3X AI citations in 30 days • Deepgram: 24X organic traffic (37K → 1.5M visitors in 60 days) Built with AI-assisted workflows. Runs on human + AI collaboration. 30-day results vs 6-month SEO cycles. Want to see how you rank in AI search? Like + comment "SEO" + repost, and I'll DM you the free scanner. (must be following)

Aryan Mahajan

28,451 views • 8 months ago

Created with Seedance 2.0 on BudgetPixel AI Prompt:Create a highly detailed, cinematic 15-second animated video set in a lively modern city completely operated by anthropomorphic animals. Use polished feature-film-quality 3D animation, expressive but natural animal movements, realistic fur simulation, vibrant environments, smooth camera motion, and consistent character design throughout. 0–3 seconds — Bear Police Officers Open with a wide establishing shot of a busy downtown intersection during a bright morning. Two large brown bears wear neat navy police uniforms with badges and caps. One bear confidently directs traffic with clear hand signals while the other stands beside a police car with flashing lights. Animal pedestrians cross the street naturally in the background. 3–6 seconds — Penguin Ice Cream Shop Use a smooth whip-pan transition to a colorful vintage ice cream truck. Two cheerful penguins wearing small aprons and bow ties serve ice cream cones through the window. A rabbit customer receives a tall strawberry cone while other animals wait in line. Include playful expressions, realistic melting ice cream, and small natural movements. 6–10 seconds — Monkey Bus Driver Transition as the ice cream truck passes in front of the camera, revealing a city bus. A friendly monkey in a professional driver’s uniform drives confidently through the busy street. Rabbits, deer, foxes, and pandas sit inside as passengers. Show the monkey checking the mirror, turning the steering wheel, and stopping smoothly at a bus stop. 10–13 seconds — Busy Animal City The camera follows the bus through the city, revealing raccoons cleaning the sidewalks, squirrels selling fruit at a street market, giraffes working near tall buildings, and birds delivering letters between rooftops. The city should feel organized, energetic, and full of believable activity. 13–15 seconds — Grand Final Shot End with a fast cinematic crane shot rising above the central city square, showing hundreds of animals working, shopping, driving, and socializing together. A fountain sits in the center while the animal city stretches into the distance under warm golden sunlight. Style: premium cinematic 3D animation, playful family-friendly comedy, highly detailed fur and clothing, expressive faces, natural body movement, realistic lighting and shadows, colorful urban production design, smooth transitions, dynamic tracking shots, shallow depth of field, 4K, 24fps, widescreen composition. Avoid: character duplication, changing uniforms, distorted paws, extra limbs, floating objects, unreadable signs, unnatural walking, chaotic traffic, stiff animation, low-detail backgrounds, sudden scene changes, text overlays, logos, subtitles, or watermarks.

Sarah Parker

58,416 views • 4 days ago

20 days ago, I connected Claude Code to my newly created instagram handle.. I gained 4.3M views and 6500+ followers in less than a month [ i post Ai generated animated stories ] Full workflow: i let claude study my account before i write another reel.. This is the cleanest content workflow i've built on claude. give it your IG first. 4 prompts handle the rest.. niche research, the reel script, the hook, and the daily automation.. the whole loop is basically, give claude your IG → find what's working → write retention-optimized scripts → engineer the hook → automate the daily output.. ▫️ Setup: give claude your instagram open claude code. claude code has a built-in web tool that browses any public URL. or install any agentic browser like Browser Harness or Firecrawl or Comet browser paste this with your handle filled in: "Browse and pull the last 30 reels and posts. Analyze my recurring topics, top-performing hooks, formats, and engagement patterns. Then map out my actual audience and what they consistently respond to." claude reads your profile, pulls every reel down, and now has the context to personalize every prompt below to YOUR account, not a generic niche. if you're on claude desktop, the same works with firecrawl MCP connected. ▫️ Prompt 1 find what actually goes viral in your niche: "Analyze the highest-performing Instagram Reels, TikToks, and Reddit posts in the [niche] niche from the last 30 days. Identify repeating hooks, visual styles, emotional triggers, and content formats that consistently generate high engagement. Then summarize the 5 strongest content angles optimized for AI-generated content and short-form videos." run this after the setup. you get 5 angles backed by what's already working in your niche, cross-checked against what's already working on YOUR account. ▫️ Prompt 2 write a high-retention reel script "Write a short-form Instagram Reel script about [topic] with an aggressive hook in the first 2 seconds. Create immediate curiosity, tension, or controversy to stop scrolling, then deliver a fast and satisfying payoff. Keep it under 30 seconds and optimize the structure for watch time, replays, comments, and shares. Finish with a subtle CTA." the line that matters: "optimize the structure for watch time, replays, comments, and shares." claude writes for the metrics, not just the word count. ▫️ Prompt 3 engineer better hooks "Study the top-performing Reels in [niche] and break down the hook structure, pacing, and emotional triggers used in the first 3 seconds. Then generate 5 new hook variations that are even more curiosity-driven, emotionally charged, and optimized to stop scrolling instantly. Focus on triggers like surprise, fear, ego, urgency, or desire." most reels die in the first 2 seconds. this prompt has claude reverse-engineer what already works, then give you 5 sharper versions to swap in. ▫️ Prompt 4 automate the whole workflow "Build a complete AI-powered content workflow for Instagram in the [niche] niche. The system should identify trending topics daily, generate high-retention scripts, create matching AI visuals, turn them into short-form videos, and generate optimized captions and hashtags. Structure everything as a repeatable workflow designed for consistent daily posting and growth." once the niche and script structure are validated, this turns it into a daily loop. one prompt that handles topic → script → visual → video → caption. these 4 prompts are the building blocks. the setup is what makes them yours. your real value is in the [niche] you plug in. content workflow built in one weekend, daily posting on autopilot from monday.

Axel Bitblaze 🪓

199,569 views • 1 month ago

$KNDX 🤖 Theres 3 big narratives that are sending coins left right and centre rn. 🚀 #AI, #Gamefi, & #NFTs 🔹Theres 50% mindshare for #AI. 🤖 🔹#GameFi mcap is hitting ATH's with #OfftheGrid, $XBG and $SUPER making spectacular moves. 🎮 🔹NFTs and the #Metaverse are making a strong comeback with $APE up 100% over the weekend. 🐵 What if there's a project that touches all these trending narratives with groundbreaking technology to disrupt all 3 of them? 🔥 💡- That's where $KNDX comes in. -💡 Kondux is a cutting-edge Web3 SaaS platform, combining NVIDIA’s Omniverse, AI, Blockchain, and dynamic NFTs to revolutionize secure asset management across industries. 👏 Their flagship product, kNFTs, are 3D digital assets usable across Metaverse and Gaming platforms, AR/VR/XR environments, and manufacturing applications. Kondux’s scalable model opens new revenue streams by enabling effective digital asset monetization. 💰 Kondux is the first Web3 project to integrate VFX pipelines with NVIDIA’s Omniverse and bringing it onto the Blockchain. ⛓️ It is also the only Web3 project with a *Select Status Partnership* with NVIDIA, operating under NVIDIA NDAs and working with them directly for more than 2 years. About their NVIDIA Integrations: 🤖 🔹There are three areas of the Kondux tech stack that coincide with three divisions of NVIDIA: 📡GDN (Graphics Delivery Network, the backbone of GeForce Now) 💡Omniverse for 3D aspects such as, geospatial data, real world physics, lighting, and raytracing 🤖NVIDIA AI Foundation, which covers many aspects of #AI, including inference and deployment scaling. The convergence of all these components lie within .USD file format . 🔹 They are the first blockchain project to integrate NVIDIA’s Omniverse Cloud and Graphics Delivery Network (GDN) to provide high-quality 3D content accessible on any device without requiring high-end hardware. 🔹 This setup streamlines content management, democratises access to resource-intensive 3D content, and enables real-time interaction with 3D NFTs. Now, I haven’t seen any crypto project so deeply connected with NVIDIA and NVIDIA technology. GDN is a HUGE competitive advantage. With it, the need for #GPU’s basically goes out the window. 🤯 Now lets take a look at some of the other main features... 👀 OpenUSD (Universal Scene Description): 📽️ 🔹 Kondux is leveraging USD technology, developed by Pixar and used by Meta, Apple, Microsoft and other industry leaders to enhance 3D graphics and interoperability within its creative ecosystem. 🔹 Originally created for high-end film production, USD now supports a variety of applications, including gaming and virtual reality, making it a key asset for Kondux. kNFT's: 🎨 🔹 Kondux is pioneering a new category of NFTs known as kNFTs, which aim to redefine NFT utility through innovative features. 🔹 A standout feature is the upgradeable aspect provided by Kondux DNA, allowing kNFTs to transform and combine with other NFTs, creating limitless possibilities in art, gaming, and music. 🔹Through the Kondux AI portal it will be possible to communicate with kNFTs. They can learn and adapt. This AI technology is revolutionary because it makes human to kNFT interaction possible, turning it into a unique, personalized experience. Check out the clip of kNFTs in Unreal Engine 5 gameplay below. 👇 Kondux is a very obvious utility play with huge upside because it’s multi narrative. 📈 It's seriously groundbreaking stuff that they’re about to launch. 🚀 After speaking with the team there’s no doubt in my mind this will do crazy big numbers in the next months. 🤑

Altcoin Miyagi🇯🇵

17,303 views • 1 year ago

Season 2 of Valiants: Tap-Tap is here!🌟 It’s time to dive into the new and exciting world of Valoria! Here’s everything you need to know: Now available here! 🔥 New Points Store Your efforts in Season 1 have paid off. Now, you can use your store points to redeem amazing rewards. What will you find? Surprise boxes containing exclusive items in Valiants: Tap-Tap and Valiants: Arena, NFTs that you can trade in the future, and a portion of the $VGN Airdrop allocation. Important: For now, the boxes can only be purchased, but soon you’ll be able to open them to discover what’s inside. Stay tuned for updates so you don’t miss anything! 🕒 🎁 Wild Spin - Spin and Win Introducing the new **Wild Spin**, where each spin could change your fate. Spins will be awarded based on the activity and performance of your friends with a good User Score. So, stay active and keep connecting with your teammates to maximize your chances! What can you win? Experience, unlocks, and even USDT prizes that you can withdraw directly to your airdrop wallet. Tip: Use your USDT winnings to purchase additional spins and keep the wheel turning. How far will your luck take you? 🎰 💥New Valiants and Items The battle in Valoria heats up with the arrival of new Valiants and accessories. These exclusive items will allow you to explore new strategies and challenge your opponents like never before. 📋 Patch Notes Not only are we introducing new features, but we’ve also made several improvements and adjustments to provide you with a smoother and more satisfying gaming experience. Here are all the details: - Daily Combo Changes: The combo is now linear, meaning the challenge has increased. Can you keep up the streak? This change is designed to reward the most skilled and dedicated players. - Daily Login Update: Starting from day 10, the rewards get even better. Don’t miss a single day to maximize your benefits! We’ve also adjusted the progression so that each login brings you closer to more significant rewards. - Complete UI Redesign: We’ve revamped the entire user interface to make it more intuitive and visually appealing. Additionally, you can now choose between two visual themes: the serenity of Kai or the energy of Mimi. We’ve even added some music to accompany your adventure! These improvements are designed to make your time in Valoria even more enjoyable. We hope you love them as much as we do! ✨Don’t miss out on this new adventure! Explore all the new features in Season 2 and discover how to dominate the new world of Valoria. Remember, every decision can bring you one step closer to victory. See you in the game!🔥

Valiants

81,606 views • 1 year ago

this effect is all over tiktok right now and nobody's explaining how to actually do it properly... the 3d balloon character thing. where someone turns into a shiny inflatable version of themselves that still moves and talks. looks pretty smooth in feeds. the workflow is stupid simple once you see it. step 1: take any photo. drop it into an image gen tool (nano banana pro). prompt it with something like "make the person in the photo a plastic blow up balloon character with a shiny surface. keep the face details as 3d balloon details including the person in the background. don't change background" that's it for the image. don't overcomplicate the prompt. shorter = more consistent results. (learned this after wasting like 2 hours trying to get "perfect" prompts that kept giving me garbage) step 2: take that balloon image + your original video and drop both into kling motion control. prompt: "turn the motion and detailed mouth movement of the video to the setting of the image" that's literally it. kling maps the motion from the real video onto the balloon character. mouth moves. head turns. expressions transfer. the whole thing renders in a few minutes. the result looks like a $500 custom animation and costs you maybe $0.30 in kling credits. people are getting 500k+ views with these because the scroll-stop factor is insane. nobody expects to see a shiny inflatable version of someone giving a real speech or doing a product review. the play here is obvious btw. run this for client content (mix with the hook and real body, check the results yourself) or use it on your own faceless channels as a hook pattern before the algo catches up...

KNOX

25,773 views • 5 months ago

You can't 3D reconstruct glass from images... ...WRONG! Thanks for video diffusion, now just about anything is possible! Introducing...Diffusion Knows Transparency (DKT) Transparent and reflective objects usually break robot vision and photogrammetry pipelines because they don't follow the "solid object" rules standard cameras expect. DKT is a new AI model that repurposes the "internal physics engine" found in video generation models to solve this problem. Researchers took a massive video diffusion model (WAN) and fine-tuned it using a custom-built synthetic dataset to turn it into a high-precision depth sensor. To train the AI, they built the first massive synthetic video library of transparent objects, 1.32 million frames of perfectly labeled glass and metal objects in motion. Without ever seeing a "real" labeled video of glass during training, the model (DKT) outperformed all previous specialized systems on real-world benchmarks (ClearPose, DREDS). They created a "lightweight" 1.3B parameter version that runs fast enough (0.17s per frame) to be used on actual robot hardware. Two reasons I find this project important: 1. It further proves that synthetic data will be essential for training the next generation vision models. 2. In real-world robotic tests, using DKT's depth maps nearly doubled the success rate of robot arms trying to pick up objects on tricky reflective or translucent surfaces. At home robots will need to interact with these types of objects on a daily basis. Check out the project page here: Code is LIVE! #Computervision #Robotics #AI

Jonathan Stephens

17,712 views • 6 months ago

Remember when we as football fans had to rely solely on paper draft guides, sports radio rumors, and gut feelings to predict draft day decisions? Excited that fans now have access to the NFL's Draft IQ powered by Amazon Web Services ( – the most sophisticated tool yet for following the NFL draft and your favorite team's strategy. Draft IQ is built on Amazon QuickSight, our cloud business intelligence service that makes it easy to analyze and visualize massive amounts of data. QuickSight processes real-time data to give fans unprecedented insight into team decision-making, updating the entire draft landscape every five minutes. You can explore team needs, draft capital, and front office tendencies through personalized team dashboards, plus get AWS-powered machine learning predictions about potential trades and picks. During draft week, fans can track picks, prospects, and Next Gen Stats in real-time. We're also introducing Amazon Q Business integration, our generative AI-powered assistant. Q Business leverages large language models to understand and respond to natural language queries, allowing fans to ask detailed questions about draft prospects, team strategies, and historical draft data. It can provide AI-generated insights based on the same historical Next Gen Stats research data that powers Draft IQ, giving fans a new way to engage with the draft experience (check out the example below). Can't wait to see what stories the data tells us as teams make their selections and excited to dig into the Giants' data myself :)

Andy Jassy

102,869 views • 1 year ago

Claude Code + Google Stitch 2.0 is f*cking cracked 🤯 Google just dropped a free AI design agent that solves Claude Code's biggest weakness: frontend design. One screenshot of a high-converting landing page → a production-ready site for your brand in minutes. All inside Google Stitch + Claude Code. Perfect for DTC brands and agencies who are building advertorial pages and product launch pages for Meta but burning days on designer back-and-forth. If you're running Meta ads and need 5-10 different landing pages testing different hooks, angles, and offers — each one targeting a different audience and pain point — you know the bottleneck isn't the ads. It's the pages. Briefing designers, waiting for revisions, paying $2-5K per page. Stitch eliminates the design bottleneck: → Find a high-converting advertorial that's scaling on Meta → Screenshot it and drop it into Stitch (powered by Gemini 3.1) → Stitch redesigns it with your brand's colors, fonts, and imagery using Nano Banana 2 → Edit sections visually — headlines, CTAs, layouts — without touching code → Export the code and paste it into Claude Code → Claude builds the full production site and deploys to Vercel or Netlify in 60 seconds No designer. No $3K per landing page. No Claude Code frontend that looks like a template from 2019. What you get: → Designer-quality landing pages and advertorials built in minutes, not weeks → Visual editing so you actually see the design before you code it → Nano Banana 2 generating on-brand product imagery and hero shots → A repeatable system — new angle, new page, same pipeline Built 100% with Google Stitch 2.0 + Claude Code. I put together a full playbook showing the exact workflow: how to find winning pages, redesign them in Stitch, and deploy with Claude Code. Want it for free? > Like this post > Comment "STITCH" And I'll send it over (must be following so I can DM)

Mike Futia

125,762 views • 4 months ago