Загрузка видео...

Не удалось загрузить видео

На главную

this is pure f*cking treasure 15 GitHub projects with 1.21M combined stars that can form a real agent stack specs. memory. web data. documents. context. sandboxes. monitoring. video 01 hermes-agent ▸ 02 OpenSpec ▸ 03 caveman ▸ 04 Scrapling ▸ 05 Docling ▸ 06 PageIndex ▸ 07 mem0 ▸...

94,795 просмотров • 23 часов назад •via X (Twitter)

Комментарии: 31

Фото профиля Olle
Olle20 часов назад

cool list but holy visual larp slop 😭

Фото профиля kaize
kaize20 часов назад

how many of them do you use on a daily basis? and which one do you think is the best?

Фото профиля carp
carp20 часов назад

docling is the sleeper here.

Фото профиля Hussain Hashim | Building SundayBack
Hussain Hashim | Building SundayBack21 часов назад

@beamnxw the memory part is such a game changer. while building my stuff, I realized it’s crucial for real context handling.

Фото профиля whemo
whemo21 часов назад

gh is such an amazing place, honestly don’t know what isn’t on there

Фото профиля Slonski
Slonski23 часов назад

i already have hermes the rest of the list i'll leave

Фото профиля kozh ./
kozh ./23 часов назад

It's worth getting this kind of good stuff every night and paying for it 😄

Фото профиля 0xbobaa
0xbobaa21 часов назад

Wow, these are some really cool repo

Фото профиля ISOfunds
ISOfunds21 часов назад

the loop in 8 verbs: define, collect, parse, save, compress, run, watch, ship. harness without a vendor

Фото профиля barnyx
barnyx23 часов назад

the fact that mem0 and Fabric show up here is huge. context compression is what actually determines if agents stay useful at scale

Фото профиля broke boy
broke boy22 часов назад

15 gems on the way

Фото профиля Shtander
Shtander23 часов назад

Bookmarking this, thanks for compiling

Фото профиля expemilly
expemilly18 часов назад

I'm amazed how you constantly research them?

Фото профиля monokern
monokern22 часов назад

saved this list

Фото профиля Gipp 🦅
Gipp 🦅23 часов назад

mem0 always boosted my context recall speed

Фото профиля Diam
Diam22 часов назад

There goes the weekend.

Фото профиля NO1ennn
NO1ennn21 часов назад

I save it absolutely banger

Фото профиля TTD 🇮🇩
TTD 🇮🇩17 часов назад

Agreed

Фото профиля AI News Daily
AI News Daily22 часов назад

This is a useful way to think about agents as a stack instead of a single magic model. The part I would add is failure visibility. Memory, sandboxes, and monitoring only help if the system makes its uncertainty easy to inspect. Which layer do you think is still weakest?

Фото профиля Egor
Egor23 часов назад

open-source components still need orchestration to work as one stack

Фото профиля Imhoxbt
Imhoxbt21 часов назад

The loop is the stack. The repos are just the parts

Фото профиля shikamaru
shikamaru23 часов назад

saved for tomorrow :)

Фото профиля airplanestar 𓂀
airplanestar 𓂀22 часов назад

fair agree, curating open source primitives into a complete agent stack saves months of stitching things together from scratch, letting developers focus on shipping actual features instead of reinventing plumbing

Фото профиля BestAIprice | The World’s Cheapest AI Tokens
BestAIprice | The World’s Cheapest AI Tokens19 часов назад

Love the stack. Add a budget per run early - retry loops are where “cheap” agents get expensive 😅

Фото профиля Secta
Secta23 часов назад

mapping the workflow gives each project a defined role in composition.

Фото профиля Mika
Mika22 часов назад

i have been looking for something like this for a long time thank you

Фото профиля twinedon
twinedon23 часов назад

docling and mem0 alone make this list worth it

Фото профиля laura.st
laura.st20 часов назад

@threadreaderapp unroll

Фото профиля Edwin | AI Systems
Edwin | AI Systems22 часов назад

Specs and skills beat vibes. Put the contract in the repo before the agent touches code. Curious what you version alongside the skill so it does not rot.

Фото профиля bccx777
bccx77718 часов назад

agents very nice cooking

Фото профиля Mikadzyki🌙
Mikadzyki🌙23 часов назад

top repos

Похожие видео

FIVE LAYERS OF AGENT ENGINEERING, EACH ONE WRAPS THE ONE BELOW IT. IF YOU SKIP LAYER 2, YOUR LAYER 5 WILL LOOK BROKEN WHEN IT IS ACTUALLY JUST STANDING ON NOTHING. for weeks i debated harness vs loop vs graph like they were competing choices. then a stack diagram made the shape obvious. they are not choices. they are floors. 01 | prompt engineering. the message. unit of work: one input. inputs are role, instructions, examples, format. output is a single raw response. 02 | context engineering. the memory. unit of work: what stays in the window. a curator selects, compresses, and drops from query, docs, memory, prior turns, and tool outputs before the prompt runs. 03 | harness engineering. the machine. unit of work: the machine itself. gather (context + prompt) → LLM → tools or sub-agents → verifier → final response. the article calls this the operating environment. 04 | loop engineering. the system. unit of work: the run. goal + success criteria + max iterations + budget + completion check wrap around one harness pass. failed pass appends results to context and retries. 05 | graph engineering. the topology. unit of work: the graph run. goal + nodes + edges + state schema. graph routes to agent nodes, tool nodes, or human approval. a reviewer node with a different model and fresh context checks the final answer. the wrapping is the whole point. layer 5 assumes layer 4 works. layer 4 assumes layer 3 works. skip layer 2 and layer 3's verifier keeps failing without a clear reason. this is why swapping the model is a one-day project and swapping the stack is a quarter. the model is the commodity. the five layers around it are the engineering. full three-layer breakdown of the top of the stack (harness, loop, graph) in the post below.

kocer

30,675 просмотров • 19 дней назад

How to build long-horizon AI agents: behavior specs, ontologies, process supervision - my conversation with Mitchell Troyanovsky, co-founder of Basis 01:09 Why Everyone at Basis Was Whispering to AI when Stephanie Palazzolo walked in 04:12 Accounting as "an Intelligence Over the Economy" 06:11 What Makes an Agent Truly Long-Horizon 08:24 Inside an Autonomous, Multi-Day Tax Return 10:19 Agents That Hand Off Like Senior Engineers 11:17 A Brief History of Agents: From ReAct to Today 12:33 Why LLMs Have No Long-Term Memory 14:13 Why AutoGPT Didn't Live Up to Its Promise 15:51 The Three Breakthroughs: Opus 3, o1, o3 17:07 Why Reasoning Models Unlocked Agents 18:23 "Let's Verify Step by Step": The Road Not Taken 20:32 Pushing Back on the METR Chart 22:09 Why Coding Agents Won First 25:14 Why Real-World Agents Are Harder 26:55 How Accountants Verify Non-Deterministic Work 29:18 You Can't Scale Tax Returns Like Math 33:16 100 Evals Pass - So What? 35:53 Right Answer, Wrong Process 36:37 Behavior Specs, Explained 39:58 How Specific Should Behaviors Be? 42:18 Context Is Runtime Training Data 44:21 Who Judges the Judge? 46:45 The Move 37 Objection 50:02 The Magic Box Mental Model 52:41 "Nothing Has Changed Since o3" 54:56 Open-Sourcing Behavior Specs with Ankur Goyal Braintrust 59:45 Ontologies: A World for Agents to Live In 01:04:20 Documentation as Codebase 01:06:33 Why the Founding Fathers Were Context Engineers 01:09:05 Onboarding 300 Brilliant Alien Employees 01:11:10 Self-Improving Agent Systems 01:12:50 The Context Mistake Agent Builders Make 01:14:29 RL on Behavior Adherence 01:17:01 Will the Bitter Lesson Swallow the Harness 01:18:46 "Technical Moats Are Not Real Moats" 01:21:03 Advice for AI Builders

Matt Turck

22,509 просмотров • 1 месяц назад