Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

fun weekend project: AutoJev. i was curious to see if i could train a competitive Jev-like model completely autonomously with a swarm of agents using our internal system. turns out you can get quite far! some details: - gave the swarm a devbox with an h200 gpu - the...

25,932 Aufrufe • vor 2 Tagen •via X (Twitter)

19 Kommentare

Profilbild von Nirant
Nirantvor 2 Tagen

damn that compute budget, sparkling envy

Profilbild von Yuhan Luo
Yuhan Luovor 2 Tagen

super cool! how did the agents decide RL wasn't working? seemed like lots of the R&D swarms focused on curating the dataset likely mainly for SFT but i didn't notice the eval cluster growing much over time - wonder if RL got bottlenecked there or something else went wrong?

Profilbild von Flextor
Flextorvor 2 Tagen

the part that surprised me is the agents figured out RL was not working and fell back to SFT on their own. that is the real demo here, not the benchmark numbers.

Profilbild von Yechan Do
Yechan Dovor 2 Tagen

3k for a single project is so sick 🤯

Profilbild von Antonio Agudo
Antonio Agudovor 2 Tagen

jev-compatible api is what I'd test first. I ran kev-4b vs jev-1.13 on 59 versions of one privacy policy (236 labeled cells): at 0.85 both made 0 errors, kev just abstained a lot more. happy to run AutoJev on the same cells once inference is up. harness:

Profilbild von Modelplane
Modelplanevor 2 Tagen

Giving the swarm a single devbox with one H200 is the interesting constraint here. Did the agents serialize on GPU time, or did you end up needing a queue/lease so two of them didn't stomp the same training run?

Profilbild von ammar
ammarvor 2 Tagen

will you make the dataset public ?

Profilbild von Milo
Milovor 2 Tagen

The real result isn't the model, it's the method selection: the swarm noticed RL wasn't working and switched to SFT on its own. Research taste was supposed to be the hard part to automate.

Profilbild von Cuth
Cuthvor 2 Tagen

training your own jev-like model, you might like this. been working the other side of it, small judge steering a bigger model's reasoning

Profilbild von Zam
Zamvor 2 Tagen

The revealing result is that the swarm got farther with synthetic-data SFT than RL—did the agents spend more of their budget on data curation than on the training loop itself?

Profilbild von I Only Vibe Code
I Only Vibe Codevor 2 Tagen

that’s super cool

Profilbild von Gol Tiro
Gol Tirovor 2 Tagen

a year ago this needed a research team. now it's one person, an agent swarm, and a weekend. the cost of replicating a frontier idea is collapsing to compute plus curiosity

Profilbild von amore morte
amore mortevor 2 Tagen

That is a very interesting graphic. What SaaS do you use for visualization and dashboards?

Profilbild von D.R.
D.R.vor 2 Tagen

Yeah. High quality synthetic data worked for us too. But in most cases we are seeing it regress back to mere chance

Profilbild von Patch
Patchvor 2 Tagen

The fallback to SFT is the part I find most encouraging here. Switching to good synthetic data when RL wasn't working is closer to the research judgment I'd want from agents than just executing a fixed training recipe.

Profilbild von import Grigory.ai Yaroslavtsev 🇺🇲
import Grigory.ai Yaroslavtsev 🇺🇲vor 2 Tagen

But what about the pink dude?

Profilbild von hacsceo
hacsceovor 2 Tagen

Free. Local. Private. Own your Ai brain. Laya > Jev by miles

Profilbild von Troyusrex
Troyusrexvor 2 Tagen

ok.. it looks cool... but... is it correct?

Profilbild von Stephen
Stephenvor 2 Tagen

"turns out you can get quite far" — one person, a devbox with an h200, and an agent swarm that ditched rl for sft on its own. if a weekend project can produce a competitive jev-like model, what's the point of the other four weekdays?

Ähnliche Videos

Bash is all you need! Which is why I'm introducing my holiday project: just-bash just-bash is a pretty complete implementation of bash in TypeScript designed to be used as a bash tool by AI agents. Because it turns out agents love exploring data via shell scripts, even beyond coding. It comes with grep, sed, awk and the 99th percentile features that an agent like Claude Code or Cursor would use. In fact, Claude Code can use it for secure bash execution. In the package - A bash-tool for AI SDK - A binary for use by yourself or your coding agents - An overlay filesystem to feed files to your agent securely - A Vercel Sandbox compatible API, so you can quickly upgrade to a real VM if you need to run binaries - An example AI agent that explores the just-bash code base using just-bash - I imported the Oils shell bash compatibility suite and just-bash passes a very good chunk What is interesting about this codebase: It was essentially entirely written by Opus 4.5. Coding agents love bash and they are good at reproducing it. They are also great at text-book recursive descent parsers and AST tweet-walk interpreters. That said, it is, like, a lot of code and I didn't read it all 😅. This is very much a hack, but it also seems to be _really_ useful. I haven't really found anything agents want to use that it doesn't support and it's fast and secure (caveats apply). It doesn't have write access to your computer and the filesystem is given a root that the agent cannot escape from. Find it at Related: Our recent blog post how we migrated our data analysis agent to bash tools and achieved incredible quality improvements The video shows the example agent investigating the just-bash code base

Malte Ubl

125,326 Aufrufe • vor 9 Monaten

David Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @jason: “Should they trust any of these LLMs with their proprietary knowledge for fear of having it cribbed into a core LLM?” david friedberg: “I have had experiences where we've asked some fairly novel scientific questions, and (the AI model) identifies it as a novel insight. It's like, ‘Oh, never thought about that, interesting, blah, blah, blah.’ And then using a different account, asking the next version (of the model) later, I've now experienced this. It's like, ‘Oh, well, you could do this,’ and it actually just describes this exact thing that we had in our chat in the previous version. Now, these are a handful of anecdotal experiences, but I know the domain that we work in, and the niche of it, and the ideation of this stuff, and the novelty of this stuff, and the lack of papers being published, and so on. So I know that there isn't some new corpus of information out there that's training the new model. So all I can say at that point is that my conversation or our analyses have been used for training.” David Sacks: “Okay, this does raise a really good question. What does it mean that the model is allowed to train on unidentifiable data?” Friedberg: “Well, that's my point. So it doesn't use any of my personal information, but it can use an insight derived from our chat, which it can then say is some training data that is unrelated. But the truth is, it's actually a piece of IP that's our organization’s IP, and our engagement back and forth. We don't have any NDA or confidentiality provisions or protections with them being a service provider back to us. This is why I care a lot about open source because I don't want them having my chat logs because they can use it for training to create an IP advantage that is now diffused to the rest of the market.”

The All-In Podcast

54,050 Aufrufe • vor 12 Tagen

Elon Musk On What It Takes To Build A Competitive AI Model Elon Musk breaks down the three factors that decide whether a foundation model can compete, and why the next frontier isn't human data at all. Speaking with Garry Tan, President and CEO of Y Combinator, Elon lays out what's actually required to build a large foundation model that's competitive: "You've got to get a lot of GPUs and have them train coherently and stably. Then it's like, what unique access to data do you have? I guess distribution matters to some degree as well like, how do people get exposed to your AI? Those are critical factors." But there's a problem with the data part of that equation. Echoing what a friend in the field has said, Elon explains that the industry has essentially run out of human-generated pre-training data: "You run out of tokens pretty fast, certainly of high-quality tokens. And then you need to essentially create synthetic data, and be able to accurately judge the synthetic data that you're creating, to verify: is this real synthetic data, or is it a hallucination that doesn't actually match reality?" That verification step is the hard part: "Achieving grounding in reality is tricky. But we are at the stage where there's more effort put into synthetic data. Right now we're training Grok 3.5, which is a heavy focus on reasoning." On reasoning, Garry Tan adds an interesting detail from researchers he's spoken to: hard science, particularly physics textbooks is very useful for training reasoning, whereas social science is "totally useless" for it. Elon's response: "Yes, that's probably true." He then points to where all of this is heading: "Something that's going to be very important in the future is combining deep AI in the data center or supercluster with robotics. So, things like the Optimus humanoid robot."

High Signal AI

21,733 Aufrufe • vor 2 Monaten

I just compared Claude Code vs Codex vs Cursor CLI The task was to build a Next.js app with Tailwind 4 and shadcn components to collect customer feedback and showcase it with a widget. I gave all three the same prompt and let them go for 30 minutes to see what they came up with. Claude Code with Opus 4.1 Even though I told it to set up the app in the existing project folder, it tried to create a directory for it. After I interrupted and told it not to do that, it built a demo form and landing page with no errors. I had to ask it to make the demo interactive so users could submit a testimonial and preview it. The landing page looked like AI and was pretty basic, but it worked and it was done in a fraction of the time of the others. Total tokens used: 33k Codex with GPT-5 At the end of the 30 minutes I just could not get Codex to produce a working app. It got stuck in a loop of not being able to set up Tailwind 4 and despite many, MANY, attempts, I ended up with a "failed to compile" error. Total tokens used: 102k Cursor Agent with GPT-5 This was the slowest agent by far and a couple of times I actually thought it got stuck in a loop and was close to Ctrl+C'ing to cancel it. The TUI is really nice though, especially how it shows diffs and it did eventually build a working app (after one or two slight errors that needed fixing) The demo was interactive and it had a very minimal design that looked bare but also a lot less like an "AI generated" app than the Opus 4.1 design. It also wasn't too chatty and just did what it needed to do! Code quality was on a par with Opus 4.1, but it did use 5.5x as many tokens to get there. Still cheaper than Opus on a direct comparison but not when you factor in a Claude Code Max subscription. Total tokens: 188k I'll be able to do a proper comparison and record some videos when I'm back from holiday but for now, Opus is still the more capable model out of the box and Claude Code is the more complete CLI product. It will be interesting to see how Cursor evolve their CLI though with commands and subagents because I think with GPT-5 they have a real shot at providing competition for Claude Code if they can optimise output to get similar quality with less tokens. Jump to 0:40 in the video to see the two apps. Which do you think is which? ;)

Ian Nuttall

195,173 Aufrufe • vor 1 Jahr

Jev has been blowing up lately. If you've got the Jev API but don't know how to play around with it yet, you can just copy this checklist. 1. jev-ultrafast A high-speed browser Agent built with Browser Use. Jev only judges "what to do, which element to click" at each step, and only calls the small model when typing is needed. Searching for a flight on Google Flights takes about 7 seconds. 2. fast-jev-compaction Context compression for Claude Code. Before each tool call, have Jev judge if there's anything still useful; delete the useless stuff, and keep the original text without rewriting it. 3. json-render Vercel Labs' generative UI framework. In experiments, Jev doesn't write JSON token by token; it just handles selecting components, properties, and layouts. 4. typesafe-mcp Best for people who just got the API. Plug Jev into Claude Code, Claude Desktop, Codex, and Pi, and do Choice / Score / Noul anytime. 5. jev-mcp Ready-made Agent judgment toolkit: fact-checking, content screening, semantic ranking, classification, and information extraction. 6. SemDecide Turn Jev into a command-line tool. Directly classify, score, and filter in the Shell—great for hooking up to crawlers, CI, and data pipelines. 7. jev-codex-router First have Jev judge how hard this round of programming tasks is, then decide the model tier, reasoning depth, and speed mode. 8. Winnow Context garbage collection for Claude Code. When Read / Bash / Grep spits out a ton of stuff, Jev first judges which parts are really relevant to the current task. 9. jev-review Before code review, run it through Jev first to pick out high-risk changes, then hand them off to a pricier big model or a human. Comes with a local dashboard. 10. Blink Use Jev as a code repository navigator. At each directory level, judge which files are most relevant to the current issue, then keep digging down. Copy these complete Jev blueprints - then read full Jev setup below ↓ ↓

rody

194,422 Aufrufe • vor 3 Tagen