Загрузка видео...

Не удалось загрузить видео

На главную

Most world models fall apart after a few seconds. Common failure modes include texture smearing, warped geometry, and scenes that no longer look real. LingBot-World 2.0 from Robbyant seems to hold 720p at 60 fps for a full hour of interaction. That’s impressive. Here is what makes that possible.

11,920 просмотров • 2 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Most AI world models can generate beautiful scenes. Keeping those scenes alive for an hour without falling apart is the real challenge. That's what caught my attention about LingBot-World 2.0 (LingBot-World-Infinity) from Robbyant Instead of chasing longer videos, it focuses on something much harder: persistent, interactive worlds that stay coherent while you explore. A few highlights: • Generates worlds from a single frame and continuously responds to live user actions through a causal world model. • Streams stable 720p at 60 FPS in real time. The team reports a continuous 60 minute stress test across 20 different scenarios with no noticeable visual degradation. • Uses a Brain-Cerebellum co-simulation framework where a VLM plans events while the video model turns them into consistent world evolution. • Pilot and Director Agents help drive character behavior and introduce new objects and events. • Open sourced with a 14B flagship model, while the paper also describes a lightweight 1.3B version for a single consumer GPU. There is also an online interactive demo. The biggest takeaway? We're moving beyond AI that generates clips. We're getting closer to AI that generates living, evolving worlds you can actually interact with. And that feels like a much bigger shift than another jump in video quality. Explore more: 💻 Github: 🤗 Weights: 🌐 Website-with videos you can use : 🎮 Try it online: #Robbyant #LingBot #WorldModel #EmbodiedAI #OpenSource #Robotics #ad

Alif Khan

84,542 просмотров • 2 месяцев назад

This World Model 'LingBot-World-Infinity (LingBot-World 2.0)' just released from Ant Group looks realy promising. It is an open causal world model with an Agentic harness. Most interactive world models hold together for a few minutes. Then textures smear and geometry warps. That's a video model, not a world — and Robbyant just drew the line at the attention mask. They released LingBot-World-Infinity (LingBot-World 2.0) — a 14B open causal video world model built on Wan2.2, trained with a Mixture of Bidirectional and Autoregressive (MoBA) attention mask, then distilled into a few-step real-time generator with no post-hoc drift filtering anywhere in the stack. Here's what's actually interesting: → Pure teacher forcing overfits — as context grows, the model leans on context instead of predicting frames. MoBA appends a bidirectional full-attention block as a regularizer → Leak-free cross-attention: AR rows attend to background prompt a_B plus chunk prompts a_≤i, lower-triangular. Bidirectional rows see one global prompt a_G → DMD runs over long self-rollout trajectories, not teacher-forced states — the student is optimized on the distribution its own errors induce → Director-Pilot harness: a VLM proposes event cards, the DiT generator renders physical dynamics. Mode B adds a SAM tracking loop for object-centric interaction → One 60-minute uninterrupted session, 20 distinct scenarios, no perceptible decay Full analysis: Paper: Model weight: GitHub Repo: Project: Ant Group Robbyant

Marktechpost AI

233,376 просмотров • 2 месяцев назад

At Avalon we are building "Real-time creating" - the ability to generate gameplay ready persistent worlds prompted from text. While others are building real-time video world models, Avalon is building real-time world generation inside a fully playable, persistent multiplayer engine. Internally running at 3840×2180 at 60 FPS. Built on Unreal Engine. Multiplayer by default. Persistent by default. Gameplay-ready by default. This is not a video latent replay. Not a simulation of interaction. It is a real 3D world with physics, logic, and authoritative multiplayer state. Avalon is trained on proprietary Avalon interaction data and powered by a hybrid system that combines language understanding, 3D model generation, procedural systems, and structured gameplay logic synthesis. Players can walk through a live world and generate environments, assets, mechanics, and entirely new gameplay modes using natural language. We accomplish this through a combination of 3D model generation, game logic generation based on our proprietary systems, and AI driven world creation. While other players are inside it. Changes persist instantly. State is synchronized in real time. Creation happens inside the world, not outside of it. Describe a biome. Spawn a civilization. Create a survival mode. Build a dungeon crawler. Launch a new game inside the world. Avalon interprets intent and integrates it directly into the live multiplayer environment. This is not a world model predicting video. This is a gameplay engine that understands language. If you can describe it, you can build it. And others can walk into it instantly.

AVALON

65,084 просмотров • 7 месяцев назад