Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

One of our autoresearch runs sat flat for 60 steps. One message got it moving again. Everyone is pushing research agents toward full autonomy, but what helps most is being able to step in when a run goes wrong, without breaking the loop. (1/6)

182,647 Aufrufe • vor 2 Monaten •via X (Twitter)

15 Kommentare

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

We made Weco autoresearch steerable: message a run mid-flight and it branch a new search, keeping everything it found. (2/6)

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Hand it an idea from a paper. "Read this arxiv link and try it." It reads the paper and runs the idea as a new branch. (3/6)

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Or start several directions at once, each with its own step budget. "Spend 10 steps trying rust, 10 trying mlx." They run in parallel from the best node. (4/6)

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Every steer becomes its own subtree. You can see where each idea came from and which steer actually moved the metric, so you know what to build on and what to drop.

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

It works right away from a Claude Code or Codex session. We also let you connect your local Claude Code straight to the dashboard, so everything is together. Get started with "pipx install weco & weco setup claude-code"

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

More details:

Profilbild von Nick Venturi
Nick Venturivor 2 Monaten

turns out bots need babysitting too

Profilbild von Han Xiao
Han Xiaovor 2 Monaten

Why can’t it find escape path by itself? Why still need human to point out arxiv?

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Good q. Because human experts are usually still more creative, think deeper, or have global view that AI don’t have in its context window. For arxiv, it’s not about let it read papers in general but “this” paper

Profilbild von Han Xiao
Han Xiaovor 2 Monaten

imo, if it stuck in local and requires a hand from human, that's smart hill-climbing but not autoresearch or long-horizon task. frontier intelligence should be able to fire well-composed search queries, find related info and escape from local optimum by itself.

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Frontier intelligence still benefits a lot from human in the loop: It can escape from local optima but will still exhaust ideas eventually, because it still have a finite and small context window. On high level, human inputs are helpful because human intelligence is not a subset of AI

Profilbild von Han Xiao
Han Xiaovor 2 Monaten

No that I agree. I understand the human inputs r valuable, but scarce like super very scarce especially when autoresearch on nontrivial question. In ur opinion, is the vision of autoresearch just be an intern and execute our thoughts of the know-unknown or should it focus on explore the unknown-unknown?

Profilbild von Zhengyao Jiang
Zhengyao Jiangvor 2 Monaten

Yes, human input is scarce. At the same time tokens& experiments are also expensive, so the goal of the product is to improve research outcomes per unit of human effort. In the longer term the vision is the latter, explores unknown unknowns endlessly. I’m sure we will get there eventually, need consistent investment on research side to get there though

Profilbild von Josh Cason
Josh Casonvor 2 Monaten

Poked around the docs. Looks very cool 👍

Profilbild von Jack Of Spades
Jack Of Spadesvor 2 Monaten

I think the world would be better if people like you were executed by the state

Ähnliche Videos

HOW TO USE AI LOOPS TO RUN YOUR BUSINESS 24/7 A lot has been written about loop engineering for building products. Almost nothing about using loops to run the business itself. That's the bigger idea. A loop is when you give an agent a goal, a way to check its own work, and permission to keep trying until it hits that goal. Build. Verify. Repeat. Stop when the condition is met. Here's what it looks like in practice: 1/SEO loop You're position 30 for a term you want. The loop runs once a month, makes changes, checks where you rank, and keeps pushing until you're on page one. This is running in production right now on Inbox Zero. 2/Ads loop You're spending $100 a day and losing money. The loop tests creative, checks profitability, kills what fails, and keeps going until the account is in the black. 3/Eval loop Your AI feature is only 88% accurate. The loop keeps adjusting the prompt and swapping the model until it passes 90%. 4/LLM visibility loop People search in ChatGPT now, not just Google. Same loop, new scoreboard. Are we the answer or not? The whole thing hinges on one thing: a metric that comes back black and white. Where do I rank? Did it hit profitability? Did the evals pass? Give an agent that scoreboard and it runs for months. Loops used to run for 30 minutes. These run for a year. Take a step, sleep, wake up next month, take another one. You're basically hiring an agency that never sleeps, gets paid in tokens instead of invoices, and undoes its own mistakes when the number goes down. Full episode on The Startup Ideas Podcast (SIP) 🧃 watch

GREG ISENBERG

83,349 Aufrufe • vor 2 Monaten

A finance professor manages $200M with AI agents, and he told everyone why: "Large language models are at the level of a fourth-year PhD student in every field" Alejandro Lopez-Lira's AI fund, Autopilot, returned 56% last year. The S&P did 16%. There are 52,000 people with money in it, and most of them just watch the machine work. What he automated is the same six-step loop every fund on earth runs: find an idea, code it, backtest it, deploy it, read the autopsy, learn from it. A quant at Two Sigma runs that loop once a month, and the salary time alone costs around $50,000 per hypothesis. All steps from this loop now fit in AI trading text box. Plain English in, executable strategy out, five-year backtest in 12 seconds, live on a broker 90 seconds after you typed the sentence. He runs $200M with AI. You can run same AI fund in two clicks, free to try: Step 6 on this loop is where everyone is stuck. Your agent has no memory. Every strategy it kills goes into a log nobody reads, and the next one starts from zero. Nobody keeps negative results. Not Citadel, not Man Group, not a single repo on GitHub. Fix that and the agent remembers every hypothesis it killed and the regime it died in. It stops burning cycles on your old mistakes. Jane Street pays 3,500 people to run this cycle and made $39.6 billion doing it. Five sixths of it is now free. Bookmark & read full map of this loop in the article below. Most people still think AI trading is out of reach for them - it isn't. Don't want to spend a dollar for testing this? Kalshi just opened a perps exchange and gives US users $25 free to start ->

cvxv666

82,936 Aufrufe • vor 1 Monat