Loading video...

Video Failed to Load

Go Home

Got five papers accepted by #ECCV2024 European Conference on Computer Vision #ECCV2026 ! Huge thanks to all my collaborators! 😃 See you in Milan 🇮🇹 Summary of Selected Works (I made a fast-forward for them 😄) - [Shape Generation] Surf-D: Generating High-Quality Surfaces of Arbitrary Topologies Using Diffusion Models,...

18,187 views • 2 years ago •via X (Twitter)

10 Comments

Jingbo Wang's profile picture
Jingbo Wang2 years ago

@eccvconf 🥳🥳🥳

Zhiyang (Frank) Dou's profile picture
Zhiyang (Frank) Dou2 years ago

@eccvconf 🛫

Chen Wang's profile picture
Chen Wang2 years ago

@eccvconf Congrats!

Zhiyang (Frank) Dou's profile picture
Zhiyang (Frank) Dou2 years ago

@eccvconf Thanks,Chen!

Jiye Lee's profile picture
Jiye Lee2 years ago

@eccvconf Super! Congrats🎉🎉

Zhiyang (Frank) Dou's profile picture
Zhiyang (Frank) Dou2 years ago

@eccvconf Thanks, Jiye!

Heming Zhu's profile picture
Heming Zhu2 years ago

@eccvconf congrats 🥳

Zhiyang (Frank) Dou's profile picture
Zhiyang (Frank) Dou2 years ago

@eccvconf Thanks, Heming!

Ling-Hao (Evan) CHEN's profile picture
Ling-Hao (Evan) CHEN2 years ago

@eccvconf amazing works

Zhiyang (Frank) Dou's profile picture
Zhiyang (Frank) Dou2 years ago

@eccvconf Thanks, Ling-Hao!

Related Videos

DisCo: Disentangled Control for Referring Human Dance Generation in Real World paper page: Generative AI has made significant strides in computer vision, particularly in image/video synthesis conditioned on text descriptions. Despite the advancements, it remains challenging especially in the generation of human-centric content such as dance synthesis. Existing dance synthesis methods struggle with the gap between synthesized content and real-world dance scenarios. In this paper, we define a new problem setting: Referring Human Dance Generation, which focuses on real-world dance scenarios with three important properties: (i) Faithfulness: the synthesis should retain the appearance of both human subject foreground and background from the reference image, and precisely follow the target pose; (ii) Generalizability: the model should generalize to unseen human subjects, backgrounds, and poses; (iii) Compositionality: it should allow for composition of seen/unseen subjects, backgrounds, and poses from different sources. To address these challenges, we introduce a novel approach, DISCO, which includes a novel model architecture with disentangled control to improve the faithfulness and compositionality of dance synthesis, and an effective human attribute pre-training for better generalizability to unseen humans. Extensive qualitative and quantitative results demonstrate that DISCO can generate high-quality human dance images and videos with diverse appearances and flexible motions.

AK

161,479 views • 3 years ago

Multi-Track Timeline Control for Text-Driven 3D Human Motion Generation paper page: Recent advances in generative modeling have led to promising progress on synthesizing 3D human motion from text, with methods that can generate character animations from short prompts and specified durations. However, using a single text prompt as input lacks the fine-grained control needed by animators, such as composing multiple actions and defining precise durations for parts of the motion. To address this, we introduce the new problem of timeline control for text-driven motion synthesis, which provides an intuitive, yet fine-grained, input interface for users. Instead of a single prompt, users can specify a multi-track timeline of multiple prompts organized in temporal intervals that may overlap. This enables specifying the exact timings of each action and composing multiple actions in sequence or at overlapping intervals. To generate composite animations from a multi-track timeline, we propose a new test-time denoising method. This method can be integrated with any pre-trained motion diffusion model to synthesize realistic motions that accurately reflect the timeline. At every step of denoising, our method processes each timeline interval (text prompt) individually, subsequently aggregating the predictions with consideration for the specific body parts engaged in each action. Experimental comparisons and ablations validate that our method produces realistic motions that respect the semantics and timing of given text prompts.

AK

126,595 views • 2 years ago

Tencent presents GameGen-O Open-world Video Game Generation We introduce GameGen-O, the first diffusion transformer model tailored for the generation of open-world video games. This model facilitates high-quality, open-domain generation by simulating a wide array of game engine features, such as innovative characters, dynamic environments, complex actions, and diverse events. Additionally, it provides interactive controllability, thus allowing for the gameplay simulation. The development of GameGen-O involves a comprehensive data collection and processing effort from scratch. We collect and build the first Open-World Video Game Dataset (OGameData), amassed extensive data from over a hundred of next-generation open-world games, employing a proprietary data pipeline for efficient sorting, scoring, filtering, and decoupled captioning. This robust and extensive OGameData forms the foundation of our model's training process. GameGen-O undergoes a two-stage training process, consisting of foundation model pretraining and instruction tuning. In the first phase, the model is pre-trained on the OGameData via the text-to-video and video continuation, endowing GameGen-O with the capability for open-domain video game generation. In the second phase, the pre-trained model is frozen, and we fine-tuned using a trainable InstructNet, which enables the production of subsequent frames based on multimodal structural instructions. This whole training process imparts the model with the ability to generate and interactively control content. In summary, GameGen-O represents a notable initial step forward in the realm of open-world video game generation via generative models. It underscores the potential of generative models to serve as an alternative to rendering techniques, which can efficiently combine creative generation with interactive capabilities.

AK

367,128 views • 1 year ago

🚨 SIGGRAPH Asia 2025 Paper Alert 🚨 ➡️Paper Title: WorldExplorer: Towards Generating Fully Navigable 3D Scenes 🌟Few pointers from the paper 🎯Generating 3D worlds from text is a highly anticipated goal in computer vision. Existing works are limited by the degree of exploration they allow inside of a scene, i.e., produce stretched-out and noisy artifacts when moving beyond central or panoramic perspectives. 🎯 To this end, authors of this paper proposed “WorldExplorer”, a novel method based on autoregressive video trajectory generation, which builds fully navigable 3D scenes with consistent visual quality across a wide range of viewpoints. 🎯They initialize their scenes by creating multi-view consistent images corresponding to a 360 degree panorama. 🎯Then, they expanded it by leveraging video diffusion models in an iterative scene generation pipeline. 🎯Concretely, they generated multiple videos along short, pre-defined trajectories, that explore the scene in depth, including motion around objects. 🎯Their novel scene memory conditions each video on the most relevant prior views, while a collision-detection mechanism prevents degenerate results, like moving into objects. 🎯Finally,they fuse all generated views into a unified 3D representation via 3D Gaussian Splatting optimization. 🎯Compared to prior approaches, WorldExplorer produces high-quality scenes that remain stable under large camera motion, enabling for the first time realistic and unrestricted exploration. 🎯They believe this marks a significant step toward generating immersive and truly explorable virtual 3D environments. 🏢Organization: TU München 🧙Paper Authors: Manuel-Andreas Schneider, Lukas Höllein , Matthias Niessner 📝 Read the Full Paper here: 🗂️ Project Page: 🧑‍💻 Code: 🎥 Be sure to watch the attached Technical Summary Video - Sound on 🔊🔊 Find this Valuable 💎 ? ♻️QT and teach your network something new Follow me 👣, naveen manwani , for the latest updates on Tech and AI-related news, insightful research papers, and exciting announcements. #SIGGRAPHAsia2025

naveen manwani

10,578 views • 10 months ago

I’m thrilled to announce that we just released GraspGen, a multi-year project we have been cooking at NVIDIA Robotics 🚀 GraspGen: A Diffusion-Based Framework for 6-DOF Grasping Grasping is a foundational challenge in robotics 🤖 — whether for industrial picking or general-purpose humanoids. VLA + real data collection is all the rage now but is expensive and scales poorly for this task. For every new gripper and/or scene, you’ll have to recollect the dataset in this paradigm for the best perf. 💡Key Idea: Since grasping is such a well-defined task in simulation - why can’t we just scale synthetic data generation and train a generative model for grasping? By embracing modularity and standardized grasp formats, we can make this a turnkey technology that works zero-shot for multiple settings. GraspGen is a modular framework for diffusion-based 6-DOF grasp generation that scales across embodiment types, observability conditions, clutter, task complexity. Key Features: ✅ Multi-embodiment support: suction, parallel-jaw, and multi-fingered grippers ✅ Generalization to partial + complete 3D point clouds ✅ Generalization to single-objects + cluttered scenes ✅ Modular design uses other robotics modules and foundation models (SAM2, cuRobo, FoundationStereo, FoundationPose). This allows GraspGen to focus on only one thing - grasp generation ✅ Training recipe: grasp discriminator is trained with On-Generator data from the diffusion model - so that it learns to correct the mistakes (if any) of the diffusion generator ✅ Real-time performance (~20 Hz) before any GPU acceleration; low memory footprint 📊 Results: • SOTA on the FetchBench [Han et al. CoRL 2024] benchmark • Zero-shot sim-to-real transfer on unknown objects and cluttered scenes • Dataset of 53M simulated grasps across 8K objects from Objaverse 📄 arXiv: 🌐 Website: 💻 Code: A huge thank you to everyone involved in this journey — excited to see what the community builds on top of it! Joint work with Clemens Eppner , Balakumar Sundaralingam , Yu-Wei, Jun Yamada Wentao Yuan and other collaborators #robotics #diffusionmodels #physicalAI #simtoreal

Adithya Murali

24,106 views • 1 year ago

Here are 10 AI video editor GitHub repos worth bookmarking: 1. Shotcut Most actively maintained open source video editor in 2026. 14K stars. Cross-platform with AI-assisted features. Just shipped a new release April 30, 2026. 2. Kdenlive The closest open source alternative to Adobe Premiere Pro. Multi-track editing, proxy editing, VST audio, and customizable workspace. Best for professional workflows. 3. OpenShot The easiest entry point for beginners. Drag and drop, 400+ transitions, 3D titles, and AI-assisted trimming. 5,700 stars. 4. Blender Not just 3D. Blender's video sequence editor and compositing pipeline is used in professional film production. 18,300 stars. Unmatched for VFX. 5. Recordly Screen recorder with auto-zoom, cursor polish, webcam overlays, and styled frames built in. Built for demo videos and walkthroughs. 6. Wan2.1 Alibaba's open source text-to-video model. Cinema-grade 1080p generation. Apache 2.0. The gold standard for open source video generation in 2026. 7. HunyuanVideo Tencent's 13B parameter open source video model. 11.9K stars. Handles 720p and 1080p with high temporal coherence. 8. CogVideoX Apache 2.0 licensed. Loads natively via Hugging Face Diffusers. Strong prompt following and smooth frame transitions. Needs 16GB VRAM minimum. 12.5K stars. 9. Open-Sora Most starred open source video generation project at 24K stars. Full training pipeline for $200K. Production-level output quality. 10. Mochi 1 Focused entirely on motion quality. The most natural-looking physics of any open source video model. Water, fabric, and human gestures without AI jitter. Apache 2.0.

Kanika

17,726 views • 2 months ago

When I was 8 years old, my favorite thing in the world was making mixtapes for my Walkman from my big sister’s CDs. That evolved into Lego stop-motion videos, and later, I fell down the rabbit hole of editing Naruto anime clips to Evanescence & Linkin Park beats. Looking back, it’s obvious: Sound and Video were always the main drivers in my life. I think that obsession also led me to my true love: Motion Design. A huge shoutout to Kevin who showed me +10 years ago the world of Motion Design. And honestly, look at where we are now. Who would have thought that there would be such a massive hype around Motion Design and UI promo videos? From giants like Airbnb to the new generation of creators like sutoxoriginals. The visual standard in 2026 is insane. But there was always one part of the process that felt like a grind to me: Sound Design. I love the result, but I hated the process. I was spending 4-5 hours per project just searching for the right sound effects. "Key typing effect" "Mouse click" "Subtle Whoosh" and so on.. So, I decided to fix it. I sat down with a Dev friend and we spent the last few weekends building a native AI tool right inside After Effects. The concept is simple: SoundDesigner AI watches your video. It understands the vibe and exactly what’s happening and when its happening. It sees a clock ticking or a text appearing, and it generates the perfect sound for it in seconds, synced to the action in your timeline. I’ve been using it on my own projects for a few days now, and it’s wild. My sound design workflow went actually from 5 hours to 20 minutes. We are launching this in February (fingers crossed for the approval at aescripts+aeplugins) If you want to see where this goes and try it out yourself, drop me a DM

Markus Gavrilov

20,299 views • 6 months ago

NEW RESEARCH: You can now create a new robot optimized for any given task! I love this new project by Huy Ha, Shuran Song, and others. Called "Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design", it generates a robot's physical design and its controller together from a task spec. DEFINITIONS: - Reward function: A scoring rule that assigns a number to how well a behavior achieves the task. Here, it is the objective the generated design is pushed to maximize (e.g., track the target motion with low error). - Tokenizing: dividing continuous or structured data (a robot's links, joints, motor specs, states, actions) into a discrete vocabulary of symbols a transformer can process, the same step that turned pixels and audio into "language" for these models. - Diffusion transformer (DiT): A transformer trained to turn random noise into structured output through iterative denoising. Here, it generates robot bodies and trajectories instead of images. - MuJoCo: The standard fast physics simulator for robotics research (DeepMind-maintained). The Menagerie is its curated zoo of ready-to-use robot models. - CMA-ES: Covariance Matrix Adaptation Evolution Strategy, the workhorse black-box optimizer: it evolves a population of candidate designs, keeps the best, and needs thousands of simulator rollouts. - Bimanual multi-trajectory optimization: Finding one design/controller that performs well across several target motions for a two-armed robot at once, harder than optimizing for a single arm and a single motion. - BERT/MAE masked-modeling trick: Train one model to fill in whatever parts of the input you hide (words for BERT, image patches for MAE); at inference, choosing what to mask chooses the task, so masking the body makes it a designer and masking the actions makes it a controller. In practice, you give it a target end-effector motion and a reward function, and it outputs a complete embodiment (link, joint, motor, and inertial property), as well as a controller to drive it. It works by tokenizing both the body (links/joints/motors) and the dynamics (states/actions) into a compact scheme called RoboTokens, training a diffusion transformer (DiT) over them. The same model predicts dynamics using those predictions ("Dynamics Self-Guidance") to push generated designs toward higher reward at inference time. Masking different token types (using the BERT/MAE masked-modeling trick) lets the one model do three jobs: generate an embodiment, control an arbitrary embodiment, or design one conditioned on a motion. It is trained on 11 robots from the MuJoCo Menagerie (0.65 kg hand to 67.5 kg quadruped, 6–35 joints), and validated in sim and on a physical ALOHA doing cloth flinging. I like the fact that this approach inverts the entire recent robotics ideas: designing a policy for a fixed robot -> designing the robot for a fixed task. Every other approach assumes the body is given and learns a controller. Transformer Transformer takes the task (target motion + reward), then generates the body and controller jointly. In practice, it is a ~180× speedup over the standard optimizer at equal-or-better quality. It reaches "CMA-ES-level quality in seconds" and finishes bimanual multi-trajectory optimization in that is worth underlining nowadays! Also worth mentioning: this is the lab behind UMI and Handroid, that I mentioned here previously! The team seems extremely creative, i love these out-of-the-box approaches. Enjoy watching the demo of robot optimization in 3D, data acquisition, then real-life testing:

Léo

25,735 views • 7 days ago

Here's a short preview of my interview with Rowan 🛡️ from Sapien, a really unique project which blends AI, Crypto, and drawing value from human contributors cc Trevor Koverko for visibility. Summary In this episode, Rowan Stone,co-founder and CEO of Sapien, shares his path from entrepreneur to building a mission-driven company at the intersection of AI and crypto. He discusses how Sapien is tackling one of the biggest challenges in AI: sourcing high-quality, human-generated data at scale. Rowan emphasizes the importance of aligning incentives from the start, building trust with contributors, and creating a system where real people help train more useful, nuanced AI models. The conversation touches on strategic partnerships, market demand, and how onboarding and education will define the future of the data economy. Takeaways — Rowan previously sold a company to Coinbase before launching Sapien. — Sapien’s goal is to monetize human understanding for AI training. — Real-world data from real people is essential for effective AI. — The need for labeled, high-quality data is growing exponentially. — Incentives and quality control are deeply integrated in Sapien’s model. — Onboarding and contributor education are critical for scale. — Sapien sees collaboration—not just competition—as a strength. — Upskilling contributors increases data quality and platform value. — Crypto-native incentives enable transparent, scalable coordination. Timeline (00:00) Introduction to Rowan Stone and His Background (02:55) The Vision Behind Sapien (06:06) Understanding AI and Data Annotation (09:01) The Role of Humans in AI Development (12:14) Sapien’s Unique Approach to Data Annotation (14:50) Partnerships and Customer Base (18:11) Quality Control and Community Involvement (21:13) On-Chain Coordination and Incentives (23:56) Demand for AI Data and Market Insights (29:51) The Future of Data Demand (30:44) Collaboration Over Competition (32:49) Revenue Generation in Crypto (35:31) The Two-Sided Market of Sapien (39:08) Customer Success Stories (44:35) The Role of Skills in Data Contribution (47:59) The Importance of Education and Onboarding (49:20) Inspiration and Influences (50:11) Overrated Trends in AI and Crypto (51:11) Distribution Channels for Onboarding (54:19) The Impact of TikTok on User Acquisition

papiofficial ᛤ

29,967 views • 1 year ago