Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

"AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos" VLM agent+optimization to reconstruct an object geometry+kinematic structure from RGB video. Pose+articulation through time including occlusions, transparent objects. No dense pixel correspondences.

24,762 Aufrufe • vor 6 Tagen •via X (Twitter)

3 Kommentare

Profilbild von Alexandre Morgand
Alexandre Morgandvor 6 Tagen

, @n_karaev, Matthew Chang, @JitendraMalikCV, @notmahi @amazon FAR (Frontier AI and Robotics) Project page: Paper: Code:

Profilbild von Alexandre Morgand
Alexandre Morgandvor 6 Tagen

@n_karaev @JitendraMalikCV @notmahi @amazon original thread:

Profilbild von Prince Kushwaha
Prince Kushwahavor 6 Tagen

Ditching dense point tracking for an agent that scripts articulated primitives and visually verifies them solves the transparent and thin object failure mode.

Ähnliche Videos

CoDeF: Content Deformation Fields for Temporally Consistent Video Processing abs: paper page: present the content deformation field CoDeF as a new type of video representation, which consists of a canonical content field aggregating the static contents in the entire video and a temporal deformation field recording the transformations from the canonical image (i.e., rendered from the canonical content field) to each individual frame along the time axis.Given a target video, these two fields are jointly optimized to reconstruct it through a carefully tailored rendering pipeline.We advisedly introduce some regularizations into the optimization process, urging the canonical content field to inherit semantics (e.g., the object shape) from the video.With such a design, CoDeF naturally supports lifting image algorithms for video processing, in the sense that one can apply an image algorithm to the canonical image and effortlessly propagate the outcomes to the entire video with the aid of the temporal deformation field.We experimentally show that CoDeF is able to lift image-to-image translation to video-to-video translation and lift keypoint detection to keypoint tracking without any training.More importantly, thanks to our lifting strategy that deploys the algorithms on only one image, we achieve superior cross-frame consistency in processed videos compared to existing video-to-video translation approaches, and even manage to track non-rigid objects like water and smog.

AK

153,305 Aufrufe • vor 3 Jahren

INFINIT partners with Google and GoogleCloudTech to bring agentic finance to millions. Anyone can access INFINIT's AI Agents for agentic coordination in their financial apps. This partnership marks the first step towards INFINIT becoming the universal infrastructure for agentic finance. This is the foundation for agentic finance at scale. Proven in DeFi, Built for Global Finance INFINIT has proven sophisticated agent coordination in DeFi: • 559,000+ Wallets • 506,000+ DeFi Conversations • 633,000+ Agent Transactions DeFi was the start. Next is scaling these capabilities to millions of developers building the future of agentic finance. A2A Integration Unlocks Exponential Distribution INFINIT integrates with Agent2Agent (A2A), Google's open standard for AI agent interoperability. This transforms how developers access INFINIT's DeFi capabilities.​ Every application adopting A2A automatically gains access to INFINIT's agent infrastructure, exponentially expanding reach from individual partnerships to ecosystem-wide distribution. Any application can now integrate INFINIT's agentic coordination capabilities: • Wallets requiring intelligent portfolio management • Trading platforms executing cross-chain strategies • Financial services building autonomous yield optimization • Portfolio managers coordinating multi-protocol operations Developers integrate sophisticated agentic coordination in a matter of hours, while users access advanced financial strategies with agentic coordination. Google's AI Infrastructure That Enables Agentic Finance Google Cloud's Vertex AI provides the foundation enabling INFINIT's agent coordination at scale with these capabilities:​ 1. Specialization: Vertex AI's Model Garden lets INFINIT's infrastructure to automatically select the optimal LLM for each natural language query.​ 2. Personalization: Vertex AI's RAG Engine processes massive on-chain and off-chain data, enabling agents to understand user history, market conditions, and protocol details providing complete context to AI agents. 3. Accuracy: Gemini's capability feeds complete instructions to all 30+ agents across multiple blockchains without compromises resulting in zero hallucination in financial execution.​​ The Vision: From DeFi to Payments to Complete Financial Coordination This is only the beginning of INFINIT and Google's collaboration.​ Google recently launched its Agent Payments Protocol (AP2) as an extension of A2A - enabling autonomous commerce across 60+ partners including American Express, Mastercard, PayPal, Coinbase, and Revolut.​ Agents will be able to execute purchases, coordinate bookings, and manage delegated financial tasks autonomously, starting from payments. The next stage entails sophisticated agentic coordination beyond simple transactions. INFINIT provides this through A2A-compatible DeFi infrastructure where agents orchestrate: • Personalized yield optimization • Cross-chain liquidity management • Portfolio rebalancing across protocols • Multi-step strategy execution As the agentic payment ecosystem matures, INFINIT becomes the infrastructure enabling agents to not only spend capital, but strategically manage and grow it.​​ From standalone DeFi agents to the universal infrastructure for global agentic finance.​ This partnership and integration with Google and Google Cloud positions INFINIT as a key building block for agentic finance, helping shape a more transparent, efficient, and accessible financial system. The future of finance is agentic. The foundation is INFINIT.

INFINIT

177,142 Aufrufe • vor 11 Monaten

🚨 Paper Alert 🚨 ➡️Paper Title: Articulate3D: Zero-Shot Text-Driven 3D Object Posing 🌟Few pointers from the paper 🎯Authors of this paper proposed a training-free method, “Articulate3D”, to pose a 3D asset through language control. 🎯Despite advances in vision and language models, this task remains surprisingly challenging. 🎯To achieve this goal, they decomposed the problem into two steps. 🎯They modified a powerful image-generator to create target images conditioned on the input image and a text instruction. 🎯They then align the mesh to the target images through a multi-view pose optimisation step. 🎯 In detail, they introduced a self-attention rewiring mechanism (RSActrl) that decouples the source structure from pose within an image generative model, allowing it to maintain a consistent structure across varying poses. 🎯They observed that differentiable rendering is an unreliable signal for articulation optimisation; instead, they used keypoints to establish correspondences between input and target images. 🎯The effectiveness of Articulate3D is demonstrated across a diverse range of 3D objects and free-form text prompts, successfully manipulating poses while maintaining the original identity of the mesh. 🎯Quantitative evaluations and a comparative user study, in which their method was preferred over 85% of the time, confirm its superiority over existing approaches. 🏢Organization: University of Oxford , Google DeepMind 🧙Paper Authors: Oishi Deb, Anjun Hu, Ashkan Khakzar, Philip Torr, Christian Rupprecht 📝 Read the Full Paper here: 🗂️ Project Page: 🎥 Be sure to watch the attached Demo Video - Sound on 🔊🔊 Find this Valuable 💎 ? ♻️QT and teach your network something new Follow me 👣, naveen manwani , for the latest updates on Tech and AI-related news, insightful research papers, and exciting announcements.

naveen manwani

14,334 Aufrufe • vor 1 Jahr

I kept seeing #GPT-6 Astra modelling results on here and got curious enough to build the thing myself: an agentic pipeline that turns a plain RGB video into 3D assets. The living room for robots MIT CSAIL. Input: one handheld phone walkthrough of our lab kitchen (~20s). Nothing else — no depth sensor, no CAD, no asset library. ~1 day (I actually slept overnight), including human-in-the-loop. Live: It is still not perfect — thin and shiny things are still weak, a few objects are drafts, the room shell needs another pass. The loop, roughly: - Monocular video → metric scan (ViPE): camera poses + depth. This is the measuring instrument, not the output. - GPT-6 lists what should exist as separate objects, then open-vocabulary detection + tracking gives per-object masks; each object is fused and measured in metres. - For every asset, GPT-6 works in its own sandbox with tools: it writes the object as a program in a small Blender DSL (closed solids, PBR colours, hinges/drawers), builds it, renders it over the original video frames, compares it with the scan points in 3D, and iterates. - A separate GPT-6 session is the verifier — it can render the model on any frame it wants and must justify every complaint with a frame. The modeller never grades its own work. - A completeness pass renders the whole modelled scene from the video's own cameras, puts it next to the real frames, and says what's still missing. What the detector keeps missing (a row of identical cabinets) gets placed geometrically instead. - Out comes MJCF/URDF with joints. Now, adding simulation to this env.

Zhiyang (Frank) Dou

11,170 Aufrufe • vor 16 Tagen