正在加载视频...
视频加载失败
True computer use is fully general. FDM-1 uses arrow keys on a computer to steer a car in San Francisco with less than 1 hour of fine-tuning data. The action policy is critical: tuning FDM-1 to drive gets much higher accuracy than tuning just the video encoder on the... show more
72,372 次观看 • 6 个月前 •via X (Twitter)
14 条评论

FDM-1 completes complex tasks and navigates interfaces well enough to use CAD applications. Forking VMs allow us to snapshot when a successful operation completes (extrusion, selection, etc.), letting us apply test-time compute to computer use.

Inspired by VPT (Baker et al), we train an inverse dynamics model (IDM) to predict frame-by-frame computer actions. The IDM leverages 40k hours of contractor data to label 11 million hours of screen recordings—550,000x larger than the largest open-source computer use dataset.

We’ve made two main advances: the ability to train on our 11M+ hour computer action dataset and understand long-context video. Our video encoder can fit nearly two hours of 30FPS, high-resolution video into a 1M token context window, ~50x more efficient than existing SOTA.

Computer use models shouldn't learn from screenshots. We built a new foundation model that learns from video like humans do. FDM-1 can construct a gear in Blender, find software bugs, and even drive a real car through San Francisco using arrow keys.

Here’s the blog post, where you can learn more about how we trained this model:

HOLY DUCK

amazing work! any idea why the baseline for the self driving task starts out worse than random guessing though?

fascinating work on the context compression part! Wondering if the model would be able to know what is a good quality action at system 2 level. eg. when to raise/fall on playing card games 😀

Can you explain what do you mean here by “tuning the video encoder”?

nice

holy fire

> less than 1 hour of fine-tuning data

that's so cool

@comma_ai
