Загрузка видео...
Не удалось загрузить видео
Introducing EvoSkill V1, an open-source toolkit that takes a benchmark and a coding agent, and evolves it into a state-of-the-art specialist in minutes. Here’s all you need to start evolving, starting with your Claude Code ↓
30,202 просмотров • 5 месяцев назад •via X (Twitter)
Комментарии: 56

1/ How does it work? Using a pattern similar to @karpathy, EvoSkill V1 acts as autoresearch for AI agent skills. It takes a benchmark, a scoring function, and a coding agent, and first evaluates the agent on that benchmark. From there, it uses the agent's failure traces as feedback to propose, test, and iterate on its prompts and skills. Read the full technical breakdown ↓

2/ Who is EvoSkill V1 for? > Devs who want to turn their agent into a task specialist without manually writing prompts > Teams building specialized solutions for customers with custom environments > Researchers working on benchmarks, agent evaluation, or autoresearch-style self-improvement

3/ Does it actually work? Yes. We ran EvoSkill on the hardest reasoning benchmarks and the results speak for themselves: > EvoSkill pushed Claude Code to SOTA performance on OfficeQA, from 60.6% to 68.1% > SealQA benchmark increased from 26.6% to 38.7%, with no human involvement, and the same skills zero-shot transferred to BrowseComp with a 5% gain > We also had similar gains with @opencode, @OpenHandsDev, @goose_oss, and @OpenAI's Codex CLI

4/ What's next? This is the first production release from Sentient Labs' AI evolution research, where we're actively exploring how to make agents self-improve across prompts, skills, memory, and the harness itself. In short, EvoSkill V1 is just the beginning. Until then, please watch & star the repo ↓

So who else is clearing their weekend? 🙋♂️

Great Work! @SentientAGI Here is a visual map + narrated walkthrough of EvoSkill

If implemented effectively, tools like these can significantly shorten the time from idea --> demo --> production.

When AI starts to build AI, so does the pace of progress!

self-learning at it's best!

and it's only going to get better!

sick video 🕶️

We think so too ✨

cloning this tonight, gonna point it at my own lil benchmark and see what specialist pops out

Look forward to seeing it!

evolving agents against a benchmark hits different than just tuning prompts

The days of manual tuning your prompts are numbered

SPECIALISTS ON TAP

It's time to let the tap flow!

back to back bangers

we ain't done yet

Senti 🚀🚀🚀🔥🔥🔥

time to evomaxxx

wait this sounds awesome, how do i try it out

Check out our quick start guide ↓

Thanks, I installed and am testing the Evoskill !

shipping the whole toolkit in public is the move, lets anyone turn their own agent into a specialist

open-source, open toolkit 🛠️

so we teaching ai to hustle now

coooooking

designer cooked

She always does 🔥

Can non Devs use this easily?

for non devs, watch out for v2!

evomaxxxing fr

fr fr

anotha ONE

@0xsachi The alpha is on sentient mrs sentient

why claude code as the first harness tho, the file edit loop or something deeper? cursor next or not

Sentient on top

Let's go higher!

Update regarding $SENT claims. ⤵️ The snapshot window is from March 25 to April 25. View your AIIocation:

Sentient does it better, v1 can't wait for what v2 will unleash. Reminiscent of ROMA moments. Just a question, are there any elements of ROMA v2 in EvoSkill asking because of the identification of errors and finding the applicable correction immediately, similar with ROMA system

Minutes already?

Will try.

Happy prototyping!

Update regarding $SENT claims. ⤵️ The snapshot window is from March 25 to April 25. Claims for all users will be live by April 27th, 12 PM UTC. View your AIIocation:

I’m curious how people in Vietnam see this 👀

nice wins - what's the compute cost and overfit risk? auto-evolve w/o humans is spicy

Spicy indeed! No exact compute cost but cost/benefit depends on how much you’ll reuse the resulting skills.

The AI race isn't about who has the best tech — it's about who finds sustainable monetization first. Execution > narrative.

@oknextlin 🤔

bench + agent in, specialist out, that's the whole stack i needed to see

giving ai back to humanity

Time to explore EvoSkill v1 Sentient is cooking👨🍳

cook with us!

I am all in on this 🚀

