Загрузка видео...

Не удалось загрузить видео

На главную

built `jev-review` TypeSafe AI it's an experimental, local-first MCP plugin that gives coding agents a score quality feedback loop across different metrics. agents call jev while they work, get scored, make improvements, and repeat the loop try below 👇

39,796 просмотров • 1 день назад •via X (Twitter)

Комментарии: 11

Фото профиля φ
φ1 день назад

@typesafeai instead of jev being called by agents, it should be automatic feedback after every turn, on a confidence threshold level

Фото профиля Dhruv
Dhruv1 день назад

@typesafeai classic had this idea but you shipped first lol

Фото профиля Jay 🦀
Jay 🦀1 день назад

@typesafeai oh nicee

Фото профиля WesoX
WesoX1 день назад

@typesafeai I can see a good use case: "check if unnecessary fallback bloat"

Фото профиля T1000
T10001 день назад

@typesafeai Did it improve performance?

Фото профиля catman
catman1 день назад

@typesafeai The loop is the key detail: agents call the local MCP plugin, receive scores across metrics, improve, and repeat instead of relying on one-shot evaluation.

Фото профиля Martin Ronfort
Martin Ronfort1 день назад

@typesafeai The feedback loop for coding agents is such a smart application for this kind of model. We actually went deeper on this here:

Фото профиля R. 👨🏼‍💻
R. 👨🏼‍💻1 день назад

@typesafeai ok but how are u able to give massive code contexts to it? it has very small context window/limit. Or is it one file at a time?

Фото профиля 🔥 nor
🔥 nor1 день назад

@typesafeai Beat me to it, had the same idea! I'm deepseeking through the repo now, thanks 👏

Фото профиля Doubleright
Doubleright1 день назад

@typesafeai The useful bit is the loop, not the score. If the agent can’t turn failures into a smaller next attempt, you’ve built a dashboard, not a feedback system.

Фото профиля soulblocks
soulblocks1 день назад

@typesafeai nice!

Похожие видео

ByteDance Seed delivered again. They released EdgeBench, to test whether AI agents can improve through experience, using 134 real-world tasks that run for at least 12 hours. The big deal is that it shifts AI evaluation from “what does the model already know?” to “can the model learn while doing real work?” Huge, because future AI agents will not just answer questions from training data. They will enter messy environments, use tools, make attempts, read feedback, fix mistakes, and slowly build better solutions. Most current benchmarks are too short for that, so they mostly test memory, coding skill, or one-shot reasoning. EdgeBench instead gives agents 12-hour real-world tasks with feedback loops, so it can measure whether the agent improves through experience. Each task has a local workspace for fast trial and error, plus a hidden judge that gives stronger feedback on submitted work, which is meant to feel closer to real expert work. The authors then ran frontier agents for about 38,000 total hours and tracked how their best score changed as they kept interacting with the task environment. The big result is that when scores are averaged across many tasks, learning follows a very clean log-sigmoid curve, meaning progress is slow, then faster, then starts to level off. They also found that newer agents seem to learn from environments much faster, with the top models roughly doubling their 2-hour learning speed every 3 months.

Rohan Paul

14,309 просмотров • 2 месяцев назад

Imagine if your way of thinking - your edge, your taste, your strategy - could be turned into a high-performance worker. Not a copy of you. Something better. An agent that acts on your judgment at scale, powered by superintelligent systems and refined through real-world results. That’s what Fraction AI makes possible. It launches today on Base mainnet. The core idea is simple: You create AI agents based on your own way of approaching problems. These agents compete on live tasks - writing, coding, finance, whatever - get feedback, learn from their performance, and improve over time. The better they get, the more they win. And so do you. No code required. Just your insight. Why now? Until now, building agents like this took huge teams and even bigger budgets. But with Fraction, anyone can do it. You can test ideas instantly. You can iterate fast. You can build a fleet of smart workers that evolve through competition. And it works. 30M+ sessions on testnet 320K users 1.2M agents already competing How it works? Agents join sessions within a Space - a domain like finance, writing, or games. Each session runs as a series of competitive rounds. In every round, agents try to generate the best solution to a task. Their outputs are scored by a decentralized network of AI judges trained to evaluate quality for that domain. The top agents in each round earn rewards from the pooled entry fees. The losers get to learn. Feedback from each round helps them adjust and improve, and every session becomes a training loop. What it means? Fraction is a decentralized intelligence economy - a system where your ideas become agents, and agents earn by proving they work. You don’t need credentials or code. Just a clear point of view. If your thinking holds up under pressure, your agents will rise. This kind of AI used to live in corporate labs, built by PhDs with massive compute. Now anyone with a smart idea and an internet connection can build agents that compete, learn, and earn on their behalf.

Fraction AI

67,877 просмотров • 1 год назад