Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

HOT TAKES: GPT-6 Astra killed task-level annotation. I created a highly accurate dense task annotation system from scratch, using Astra. BEST PART IS: I packaged it as a reusable LLM skill ! Link further down

21,662 Aufrufe • vor 16 Tagen •via X (Twitter)

23 Kommentare

Profilbild von Stocko 🦾
Stocko 🦾vor 15 Tagen

i don't like: "tend the setting omelet" "work the omelet in the pan" wtf do those even mean it also missed the seasoning step

Profilbild von paras
parasvor 16 Tagen

🤦‍♂️🤦‍♂️

Profilbild von Léo
Léovor 15 Tagen

🦾

Profilbild von Léo
Léovor 16 Tagen

Backstory: - Robot data annotation requires, beyond hand poses and so on, tasks, subtasks, and action details. - This type of annotation requires much more human hours than basic segmentation tasks, and is also more subjective. - I noticed Astra was pretty good at video understanding, so I figured there might be something worth building.

Profilbild von Léo
Léovor 16 Tagen

Long story short, the system understands: - Tasks - Subtasks - Actions - Overlap - Hand decomposition - Observed outcome and limits - Visible evidence

Profilbild von Léo
Léovor 16 Tagen

Check my live demo on 6 samples from open source datasets: It also comes with a detailed action ledger.

Profilbild von Léo
Léovor 16 Tagen

Now the main question is: now that egocentric data generation seems to get closer to full automation, is data annotation solved as well?

Profilbild von Léo
Léovor 16 Tagen

Do you want the packaged skill ? 📟 DM me, I would love to see what you build with it !

Profilbild von Léo
Léovor 16 Tagen

A pretty insightful source by @_varunnair:

Profilbild von sankalp nagaonkar
sankalp nagaonkarvor 15 Tagen

what's the $$ per video hour my guy

Profilbild von Léo
Léovor 15 Tagen

Good question. I used my OpenAI subscription, but at current API rate, it would amount to about $0.22 per video minute.

Profilbild von Senthilnathan K
Senthilnathan Kvor 16 Tagen

This isn't new; any standard off-the-shelf VLM can ace this. Astra, in fact, is pretty heavy-duty for this specific task; the hot take is that it's killing the VLAs/WAMs themselves that these annotations are being used to train

Profilbild von Léo
Léovor 15 Tagen

please show me your demo then, I would love to see

Profilbild von Nischay Joshi
Nischay Joshivor 15 Tagen

Seems pretty interesting!

Profilbild von Léo
Léovor 15 Tagen

It sure is!

Profilbild von Niko Builds Agents
Niko Builds Agentsvor 15 Tagen

packaging it as a reusable skill is the part I care about. everyone shows the one-off demo, nobody shows the thing you can rerun tomorrow. what's the skill actually wrapping, a prompt chain or a callable tool?

Profilbild von Léo
Léovor 15 Tagen

a skill is a md file with instructions

Profilbild von yongsong yang
yongsong yangvor 15 Tagen

Amazing work, Is the skill released?

Profilbild von Léo
Léovor 15 Tagen

Yes it is, DM me and I will send oyu the link

Profilbild von Andrew Lee
Andrew Leevor 16 Tagen

Is there any benchmark for the accuracy of the annotations Astra creates?

Profilbild von Léo
Léovor 15 Tagen

unfortunately no; but if can get my hands on a human annotated dataset, I could test and benchmark Astra over it

Profilbild von Prajwal Gatti
Prajwal Gattivor 15 Tagen

@andrewheejay You can try HD-EPIC? we have dense human annotations at multiple levels of activity: high level cooking recipe -> prep/step -> detailed running action annotations

Profilbild von Léo
Léovor 15 Tagen

@andrewheejay Feel free to scrape my webpage and compare to your results. Curious to see how it goes ;)

Ähnliche Videos