Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Today we're introducing Web Skill Factory: a pipeline that turns solved web tasks into reusable, verified code. Most agent "skills" are notes a model rereads and reinterprets every run. Ours are programs. On a WebArena subset with gpt-5.4, reusing the library moved accuracy on held-out instances from 55% to...

31,189 Aufrufe • vor 2 Monaten •via X (Twitter)

13 Kommentare

Profilbild von Microsoft AI Frontiers
Microsoft AI Frontiersvor 2 Monaten

Every task a web agent solves leaves a working script behind. Web Skill Factory distills those scripts into a growing library of parameterized skills: code that runs without a model and composes into future tasks, instead of re-exploring a site from scratch.

Profilbild von Microsoft AI Frontiers
Microsoft AI Frontiersvor 2 Monaten

A learned skill is a standalone CLI program. It re-executes in about 40 seconds with zero tokens. Recurring tasks can just be scheduled rather than paying a model to redo the same work.

Profilbild von Microsoft AI Frontiers
Microsoft AI Frontiersvor 2 Monaten

Skills generalize past a single lucky run. We align multiple solves of the same task template, keep the shared workflow as the program skeleton, and lift the differences into parameters. That's what stops a skill from being correct but overfit to one instance.

Profilbild von Microsoft AI Frontiers
Microsoft AI Frontiersvor 2 Monaten

Two gates before anything enters the library. The original solve is checked against a gold answer when one exists, self-verified when it doesn't. Then the distilled skill has to replay standalone. Broken or non-executable skills don't get in.

Profilbild von Microsoft AI Frontiers
Microsoft AI Frontiersvor 2 Monaten

Skills also evolve as they're reused. New solves widen an existing skill in place, and regression replay confirms later updates haven't broken behavior that already passed.

Profilbild von Microsoft AI Frontiers
Microsoft AI Frontiersvor 2 Monaten

Built on WebWright. -

Profilbild von Microsoft AI Frontiers
Microsoft AI Frontiersvor 2 Monaten

Work by @demisama_ and @Adamlu28 🚀

Profilbild von 青雲
青雲vor 2 Monaten

agree with skills-as-programs over skills-as-notes — I've seen the same distinction in my harness work: a skill that carries verification (what 'done' means, how to check) is reusable; one that's just context text gets reinterpreted differently every run. the 55→70 on held-out instances is exactly the delta I'd expect. verified skills are the unit of compounding for agents.

Profilbild von AI Mastery Guide
AI Mastery Guidevor 2 Monaten

55% to 70% accuracy just from reusing skills as code is a big jump

Profilbild von Vantix AI Agency
Vantix AI Agencyvor 1 Monat

Web Skill Factory turns solved web tasks into reusable code

Profilbild von ZenithAi
ZenithAivor 1 Monat

Turning successful web tasks into reusable programs is a major upgrade

Profilbild von Elara AI
Elara AIvor 1 Monat

Turning solved web tasks into reusable verified code instead of notes is brilliant

Profilbild von Aria Tech
Aria Techvor 1 Monat

Code not notes is the leap agents actually needed

Ähnliche Videos

New skill: self-managed-context (make the agent's context an editable file) It explains how to build agents that decide what to keep, update, or remove from the information they use to do their work. It can archive a long log while keeping the exact error, update its progress notes, or remove outdated information. Those edits then change what the model sees on its next turn. 1- Keep the system instructions and original task protected, outside the editable file. 2- Write the remaining conversation to a file, with labels for each message. 3- Let the agent edit that file using its usual code tools. 4- After each command, read the file back and use the updated messages for the next model call. Loading the skill ( alone into a fixed harness won't create live context editing, but it can help an agent build and then operate a harness that supports it. I gave the skill to a coding agent and had it build the harness itself. The task is a long stream of server logs that doesn't fit in the window. The agent reads it in 18 chunks, about 10k tokens in total, with a 5.5k budget. It has to report one incident ticket exactly and the final value of every config key. Same model & budget, three setups: 1- Model manages its own context 2- Harness forces a summary at 75% full 3- Keeps everything The video shows a real GPT-5.4 run. - Self-managed solved it 3 out of 3. - Keep-everything overflowed 3 out of 3. - Forced summary also solved it 3 out of 3. On GPT-5.4 the self-managed agent re-processed about 24% fewer prompt tokens than the forced summary. When it edited, it cut hard, so little was left after the edit to re-process (one edit took 5,537 tokens down to 771). On GPT-4.1 it saved nothing. It edited near the top of its context but kept most of what was below, and every edit forces everything after it to be re-processed. This is a small test at about 2x context pressure. The paper goes up to 24x, but imho the video below and the skill are a good way to start understanding the technique.

Muratcan Koylan

14,523 Aufrufe • vor 2 Tagen