Xuanchi Ren's banner
Xuanchi Ren's profile picture

Xuanchi Ren

@xuanchi131,480 subscribers

Senior Research Scientist @NVIDIAAI. PhD @UofTCompSci. Working on GenAI and world models

Shorts

Back when we were developing GEN3C, we often imagined a Holodeck-like future: a simulator where multiple agents can enter the same generated world, act independently, and learn to collaborate. Gamma-World makes this feel more concrete. It is a generative multi-agent world model that takes synchronized observations and actions, then rolls out what each agent will see next in the same evolving world — action-responsive at 24 FPS. For me, the key challenge is going beyond two players. As more agents enter, identity cannot be tied to fixed slots, interaction cannot rely on dense pairwise attention, and independent actions still need to resolve into one shared state. Two ideas make this work: 1⃣ Simplex RoPE Distinct agent identities without slot bias — unique, but permutation-equivalent. 2⃣ Sparse Hub Attention Agents communicate through learnable hubs instead of dense all-to-all attention: agent → hub → agent This keeps cross-agent communication scalable. The exciting part: training on two-player data can generalize to four-player rollouts without additional training, and the same formulation extends to real-world bimanual robot coordination. A step toward populated world models: many agents, one shared world. Congrats to the team on Gamma-World! Project:

Back when we were developing GEN3C, we often imagined a Holodeck-like future: a simulator where multiple agents can enter the same generated world, act independently, and learn to collaborate. Gamma-World makes this feel more concrete. It is a generative multi-agent world model that takes synchronized observations and actions, then rolls out what each agent will see next in the same evolving world — action-responsive at 24 FPS. For me, the key challenge is going beyond two players. As more agents enter, identity cannot be tied to fixed slots, interaction cannot rely on dense pairwise attention, and independent actions still need to resolve into one shared state. Two ideas make this work: 1⃣ Simplex RoPE Distinct agent identities without slot bias — unique, but permutation-equivalent. 2⃣ Sparse Hub Attention Agents communicate through learnable hubs instead of dense all-to-all attention: agent → hub → agent This keeps cross-agent communication scalable. The exciting part: training on two-player data can generalize to four-player rollouts without additional training, and the same formulation extends to real-world bimanual robot coordination. A step toward populated world models: many agents, one shared world. Congrats to the team on Gamma-World! Project:

304,145 Aufrufe

Zero-shot video reasoning (chain-of-frames) isn’t just for Veo3 — open-source models can understand and edit too! 🕹️ ChronoEdit brings temporal reasoning to image editing. 🔗

Zero-shot video reasoning (chain-of-frames) isn’t just for Veo3 — open-source models can understand and edit too! 🕹️ ChronoEdit brings temporal reasoning to image editing. 🔗

22,661 Aufrufe

📢🚗✨ Excited to announce InfiniCube, our scalable generative model for dynamic 3D driving scene generation with high fidelity and controllability! InfiniCube generates very large-scale (300m×400m ~ 100,000m^2), dynamic 3D driving scenes given HD maps, 3D bounding boxes, and text prompts as controls. Website: Props to Yifan Lu for leading this project!!!! [1/N]

📢🚗✨ Excited to announce InfiniCube, our scalable generative model for dynamic 3D driving scene generation with high fidelity and controllability! InfiniCube generates very large-scale (300m×400m ~ 100,000m^2), dynamic 3D driving scenes given HD maps, 3D bounding boxes, and text prompts as controls. Website: Props to Yifan Lu for leading this project!!!! [1/N]

26,602 Aufrufe

Videos

Keine weiteren Inhalte verfügbar