Loading video...

Video Failed to Load

Go Home

Jev has been exploding in popularity recently. If you already have access to the Jev API but aren’t sure how to start experimenting with it, just copy this checklist: 1. agent-desktop Desktop automation. Read the system's accessibility tree, judge which button, menu, or input field to click next. 2....

201,620 views • 1 day ago •via X (Twitter)

19 Comments

rody's profile picture
rody22 hours ago

join substack to receive AI alpha, fresh model templates and copy-paste configs👇

Romain Pellerin🇫🇷🇺🇸's profile picture
Romain Pellerin🇫🇷🇺🇸20 hours ago

Nice list, will definitely try Prism and Canny. We just released Jev-enabled to triage and route prompts to the right model(s). Would love your feedback on the tool!

Deep's profile picture
Deep1 day ago

desktop automation finally made sense to me when i stopped guessing at coordinates and started reading the accessibility tree.

WAGMİ 100x💎's profile picture
WAGMİ 100x💎22 hours ago

accessibility tree > screenshots for agent reliability — cheaper, deterministic, zero vision hallucinations. real question: how far is jev from executing on-chain actions? an agent that reads a dexscreener alert and swaps autonomously is the actual endgame.

Jay Zhou's profile picture
Jay Zhou20 hours ago

the accessibility tree works until an app paints its own ui and leaves no tree to read. does agent-desktop fall back to vision there?

Nitiksh · NTXM's profile picture
Nitiksh · NTXM20 hours ago

That agent-desktop item is the real unlock on the checklist. Once the agent can read the accessibility tree and pick the next button or menu instead of guessing from pixels alone, the loop gets a lot less brittle. The failure mode I keep hitting on Windows is not "can it click". It is wrong-window clicks, focus stealing, and no clean stop when the agent is about to trash the focused app. Observe, ground to an element id when you can, then act, with an activity trail you can audit. NTXM Agent's Operator path is my early take on that (accessibility + set-of-marks, real terminal sessions, Keep/Revert on file edits). Still in development. Operator notes:

Hafiz Siddiq's profile picture
Hafiz Siddiq23 hours ago

The interesting shift is from code completion to evaluated decisions. Which Jev workflow has been most reliable in practice: UI control, game agents, or judging its own tool outputs?

ShadowAguy's profile picture
ShadowAguy23 hours ago

The coding loop diagram looks clean but the desktop automation step is doing the real heavy lifting. How much of the value is the agent writing code versus the agent actually executing on a live system?

Brian Hadu's profile picture
Brian Hadu19 hours ago

what’s been the most surprising result from using the agent-desktop so far?

拾光的 AI 应用观察's profile picture
拾光的 AI 应用观察16 hours ago

agent-desktop 那个 accessibility tree 读法的坑我踩过,按钮层级一深判断就飘,得配合视觉兜底。这个checklist倒是把入口讲清楚了👍

MINT's profile picture
MINT16 hours ago

smart, semantic, diff, test ‘grep’

Misha_Cripto's profile picture
Misha_Cripto18 hours ago

"Ten blueprints and a Mario-playing agent is the most honest demo of decision-making I've seen — nobody ships a README that fun. The Super Mario one reading RAM instead of screenshots is the real flex here."

morph's profile picture
morph1 day ago

these are so alphas repos, thank you!

安叫兽|Bird🕊️ 🔶 BNB's profile picture
安叫兽|Bird🕊️ 🔶 BNB1 day ago

刚拿到权限的人,照着这份清单试不容易迷路

Brjan | AI Builder's profile picture
Brjan | AI Builder14 hours ago

the checklist could be clearer, mentioning specific use cases would help a lot

Yiqi's profile picture
Yiqi20 hours ago

很齐全了!

Ctl-Alt-Del's profile picture
Ctl-Alt-Del20 hours ago

They are open weights; self host people.

notgwapo's profile picture
notgwapo13 hours ago

sounds like a solid start. lowkey curious what else you can unlock with that API. 👀

promptgm's profile picture
promptgm22 hours ago

copy-paste configs alone gonna save me hours of tweaking subbed send that alpha bro

Related Videos

Jev builds the MOST POWERFUL trading agents and someone JUST open sourced jev-trader, a fully working 24/7 trading bot with Jev along with COMPLETE low latency CODEBASE WHAT THIS MEANS FOR YOU - you no longer have to build a trading bot with Jev from scratch, you just clone this and make it yours here is how you make your own Jev trading bot with this repo: 1. clone it and run three commands, it boots straight into dry run mode with real book data, real decisions, and simulated fills so you can watch it think with zero capital 2. drop in your Jev API key and the model starts answering buy or sell on every block with calibrated probabilities in 81 milliseconds 3. swap the book reader for your own venue, the model interface is clean so any order book that returns bids and asks plugs straight in 4. tune the decision cadence and horizon, ask the model every N blocks about the move over the next M, so you control how aggressive the engine trades 5. the hot loop already fits one block with exactly two round trips, one to read the book, one to send the order, nothing else on the path, this is the institutional latency discipline most retail bots never reach 6. plug in the live server and every block, every decision, every fill streams to a public dashboard so you watch your engine run the whole point is this repo hands you HARDEST part for FREE - > the low latency engine the COMPLETE breakdown of how i turned this into hedge fund grade HFT trading system is in my article below:

Roan

119,605 views • 3 days ago

If you are confused about why 𝗝𝗲𝘃 is being called the "Internet" moment for the AI industry. This is 100% worth your time. In fact, you should watch it: It tells LLMs what to do next, in milliseconds & at almost zero cost. If you set it up correctly, you will have the AI engineer’s setup for 2028. How to set up & use Jev (to actually get the 100x): 1. Join the waitlist; it's fairly quick: typesafe .ai. 2. Then go to Claude Code or Codex. 3. Choose Opus 5-Low or Sol-Low. 4. Copy and paste this prompt: "[claude or codex] plugin marketplace add typesafe-ai/skills [claude or codex] plugin install typesafe@typesafe-ai" 5. When you type /typesafe, the skill shows up. 6. Paste your API key once and click "Allow" 7. Start with $5 in free credit. It's hard to spend more. ----- Now, here are the 3 ways to actually use Jev: 1. Jev for Linkedin I have 38,000 connections & invitations on LinkedIn. I have a new company to launch. I need to find a couple of hundred people to message. to do while saving time: > Export your LinkedIn connections and invitations. > Connect Claude to GitHub, Vercel & Apify. > Create an Apify API key to enrich your data. > Go to LinkedIn Settings → Data privacy. > Get a copy. LinkedIn will email you a ZIP file. > Open the file & find the Connections CSVs. > Upload the files to Jev. > Use Jev to classify your contacts. > Review the shortlist. 2. Jev for Gmail To go through all of my Gmail contacts and email the right people. > Go to Google Contacts. Open Other contacts. > Select all contacts. Export them. > Upload the file to Claude Code or Codex. > Use Jev to sort them into: Keep, Review, Remove and Review everything before removing anything. 3. You got lost in Claude Code, GitHub, Vercel, Apify, Jev, Typesafe. I feel you. It is overwhelming. That’s why I included the entire copy-and-paste prompt for each use case in the newsletter: A 45-second TL;DR by Matija Sosic.

Ruben Hassid

164,733 views • 2 days ago

Another insane Jev use case! Jev is making it dramatically cheaper to evaluate what actually happened inside an agent run. And finally, someone open-sourced a self-improving memory layer that can put that signal to work across agent harnesses: - Claude Code - Codex - Cursor - OpenCode, and 20+ more Beacon by Asymptote Labs continuously captures your agent history across harnesses and uses Jev to identify which runs are actually worth learning from. It then turns the highest-signal workflows, corrections, and debugging patterns into reusable skills. GitHub repo: (don’t forget to star it ⭐ ) Beacon preserves the complete session history. But preserving a run and learning from it are two different things. Most coding-agent sessions contain routine exploration, failed commands, and fixes that only apply to one task. The trace can remain available for inspection without turning every detail into guidance for future agents. Jev scores each run for evidence, reuse potential, and human correction signals. An application policy then decides whether to promote, review, or discard it. The recording shows this in action. Claude receives a coding task, modifies the implementation, and runs the tests. I then provide an edge-case correction, so Claude updates the code and adds regression coverage. Beacon automatically captures the complete session. Jev evaluates whether the correction contains a reusable engineering lesson. Once approved, that lesson becomes available to other coding agents working on the project. Since it works across harnesses: - Claude Code sessions can teach Codex. - Cursor debugging can improve OpenCode. So a problem solved by one agent should not need to be learned from scratch by another. If you want to dive deeper into Jev, I also wrote a hands-on guide to building this Jev-style decision path with open models, entirely locally. Read it below.

Avi Chawla

272,319 views • 2 days ago

Jev is cool. So is it's OSS companion, Laya. The Latest Cool Thing In AI™ tends to get a lot of hype, sometimes without everyone even understanding it. So... what is this thing? Jev is an AI model that consumes input and produces output VERY differently than chat, claude, grok. The input is two things: 1) Text state to assess. Email, html, code, whatever. 2) A set of questions which will be asked about the attached state. The canonical example from TypeSafe's docs is to identify the urgency of a support ticket. We pass the model the customer text + a single noul question "is this urgent?". Jev returns a full set of JSON. This JSON is not generated with token-by-token autoregression. Jev is not trained to produce sequences of text tokens, rather to answer questions, and guarantees well-formed responses. In the example below, we see it produces a 0.99 probability (on a 0-1.0 scale) that the answer is "yes." Jev supports exactly three types of questions (seconds example in video): a) Noul: 0–1 probability that the answer to a yes/no question is "yes." b) Choice: Ask question with pre-defined set of answers. Jev chooses the best and assigns probabilities to each. c) Score: Ask question with pre-defined scale of answers. Jev produces a position on the scale. Jev computes answers for all questions in parallel, making responses super fast even for many questions in a single request. This might seem like a narrow set of capabilities, but in the right contexts leads to incredible potential. It also makes for a useful API / primitive for programming, since the outputs are... *ahem*... type-safe and predictable in structure. Jev is not going to replace LLMs for writing your code, auto-generating your docs, or being at the core of an agent harness. But Jev IS incredibly cool, and will be used to build a lot of amazing tech. Hope this helps.

Ben Dicken

40,810 views • 3 days ago

Jev is HERE and this is the CLEAREST explanation of what it is and what NEW businesses it unlocks. (and at the end I'll tell you how to get Jev even if you're on the waitlist) WHAT IT IS You know how you open your inbox and have to decide what's junk, what needs a reply, and what can wait? Jev does that part. It looks at each thing and says "this is junk, I'm 94% sure." It doesn't write anything back to you. It just sorts. 1,700 emails for 18 cents, instantly. That sounds kinda trivial but the important part WHAT IT UNLOCKS My explanation of Jev sounds small until you realize HOW MANY jobs are exactly this. Someone reading a stack of applications. Someone deciding which support ticket goes to which team. Someone looking at inbound and deciding who's worth calling back. A few ideas on what it unlocks: 1/ Instant quotes that are actually instant. Every quote form on the internet says "we'll email you by end of day." Build the version that answers in under a second, for roofers, movers, insurance, legal intake. 2/ Lead scoring as a product. Every agency and service business has a contact form full of junk. Score every submission and send the real ones straight to the owner's phone. 3/ Support triage for companies with no support team. The ticket gets classified and routed before anyone opens it. 4/ Clipping tools. Pass in a transcript, get the best moments scored in three seconds. Every clipping product just got a cheaper engine. 5/ Application piles. Grants, permits, insurance claims, job apps, loan docs. Someone reads that stack one item at a time today. 6/ Marketplace matching. Someone types what they need and gets matched to the right local business instantly instead of waiting for callbacks. 7/ Browser agents that actually move FAST. That makes bulk browser work practical: pulling quotes from five carriers, filing the same form for 200 clients, checking supplier inventory in real time etc. TLDR; find an expensive queue and put Jev at the front of it. HOW TO GET IT I didn't realize you can skip the waitlist because Jev is live on the Vercel AI Gateway right now, so you can start calling it today. In this episode, we share how. Episode now live on The Startup Ideas Podcast (SIP) 🧃 (thanks to vogel for coming on and spilling the sauce today) Watch: Jev is a big deal because this is a whole new way to do AI Really cool Happy Jev day.

GREG ISENBERG

156,488 views • 4 days ago

LLMs vs. Jev, clearly explained! LLMs are great, and the ceiling is one you can watch scroll past: an LLM writes the answer one token at a time. give it a failed deploy and four decisions, and it produces a small JSON object where every token depends on the one before it. token nine cannot exist until token eight does, so four decisions that had nothing to do with each other just stood in a queue. then your code parses it, validates the shape, and retries when the shape is wrong. Jev fixes this without being a smaller or faster model: it removes the order. one turn on that deploy has to know: → whether the incident is urgent → which team owns it → whether the next command is risky → whether the task is actually done you declare the questions and the answer type upfront, and all four come back together, typed, with a probability on each. three primitives cover almost every fork in an agent: 1. **Choice** picks one of up to 255 options you define, like engineering, billing or sales. 2. **Score** places the state on an ordered scale you define, like low, medium or high risk. 3. **Noul** returns the probability that a yes-or-no condition is true. here is the sentence that resolves the whole confusion: text is a line you have to walk. an answer space is a room you see all of at once. ↳ generation: one order you cannot change, one string at the end, a shape you hope holds ↳ evaluation: no order at all, typed answers, a probability on every option Prompts → Agents → Loops → Graphs → Jev the probabilities matter more than the answer. ↳ engineering at 0.91 against billing at 0.09 is a route you can automate ↳ 0.52 against 0.46 is a coin flip wearing a label, and the label alone never told you which one you got that last one catches careful people. an LLM would have said "engineering" in a confident sentence and given you no way to know the race was that close. thresholds live in your code, one per action, scaled to what being wrong costs. it works when the options are known and the call depends on meaning. it is not for writing, code, arithmetic, or anything where question two needs the answer to question one. and the one that eats whole nights: type safety prevents malformed output, not incorrect judgment. Jev cannot return an option outside your schema, and it can still pick the wrong valid one with confidence. a schema-valid mistake refunds the wrong customer just as fast. an LLM writes new language when the answer space is open. Jev evaluates known paths when the answer space is closed. below i have quoted my full breakdown on Jev. it covers the three primitives, the parallel battery, the thresholds, and where it does not belong. save this and read it below ↓

Hanako

42,632 views • 2 days ago

BREAKING: Elon Musk has a very important message for the people of Wisconsin: "The general idea here is just to bring attention to the Wisconsin Supreme Court race, which is just going to be decided very shortly. This is about nine days away, perhaps just over a week, 10 days. So Somebody here might be countering one race for many reasons. The I mean, maybe the most consequential is that will decide the congressional districts. How congressional districts are drawn in Wisconsin, which if you know the other candidate wins, instead of justice schumel, then the Democrats will attempt to redraw the districts and cause Wisconsin to lose two Republican seats. That's, that's the that's, in my opinion, that's the most important thing, which is a big deal, given that the congressional majority is so razor thin, it could cause the house to switch to Democrat if federal redrawing takes place, and then we won't be able to get through the changes that the American people want. So that's that's really, it's really in my I'm just speaking personally here. My view this is about pursuing pursuing Democracy in America, and not having ridiculous districts drawn that effectively disenfranchised voters in Wisconsin. That's my number one issue. There's also a separate concern regarding judicial activism. Now, in this case, of course, the judge is elected, so there is a very good feedback loop for the people to decide whether they believe in a judge or don't believe in a judge, because they can vote judges in and out. On the federal level, we've got a bigger challenge, because judges are essentially appointed for life, which I think is something we should reconsider. I don't think anyone should have a job for life, and it would mean that the very worst judge out of several 100. No matter what they do, they're they still get a job for life, which makes no sense to me. I mean, if somebody does a bad job, especially if it's if a judge is extremely politically partisan, and if and really interprets the law in a way that is unbalanced. Then, and does a repeatedly, then in my view, a judge should be up for impeachment."

DogeDesigner

1,429,157 views • 1 year ago

AI AGENTS 101 (58 minute free masterclass) send this to anyone who wants to understand ai agents, claude skills, md files, how to get the most out of AI etc in plain english: 1. chat vs agents - chat models answer questions in a back and forth while agents take a goal, figure out the steps, and deliver a result 2. agents don’t stop after one response. they keep running until the task is actually finishedno babysitting required 3. everything runs on a loop. they gather context, decide what to do, take an action, then repeat until done 4. the loop is the system. they look at files, tools, and the internet. decide the next step. execute and then feed that back into the next step. over and over until completion 5. the model is just one piece. gpt, claude, gemini are the reasoning layer. the key is model + loop + tools + context 6. mcp is how agents use tools. it connects things like browser, code, apis, and your internal software. once connected, the agent decides when to use them to get the job done 7. context beats prompt all day. you don't need to write perfect prompts. load your agent with context about your business, style, and goals and then simple instructions work 8. claude.md or agents.md is the onboarding doc it tells the agent who it is, how to behave, what it knows, and what tools it can use. this gets loaded every time before it starts 9. memory.md is how it improves. agents don’t remember by default. this file stores preferences, corrections, and patterns you tell the agent to update it, and it gets better over time 10. skills + harnesses make it usable. skills are reusable tasks like writing, research, analysis the harness is the environment like claude code or openclaw that runs everything. basiclaly, different interfaces, same system underneath this episode with remy on The Startup Ideas Podcast (SIP) 🧃 was one of the clearest ways of understanding a lot of the core concepts of ai agents could be the best beginners course for ai agents 58 mins. all free. no advertisers. i just want to see you build cool stuff. im rooting for you. send to a friend watch

GREG ISENBERG

377,138 views • 6 months ago

Qullamaggie shows Exit Strategy for Long Swings “What’s your exit strategy for such long swings? How do you decide this? Okay, that’s easy. Let’s take an example, ROKU. So I bought it here, on the opening range size. It gapped up on earnings, huge volume. All the things I’m looking for. You know, long range break had like a 6, 7-month range break had surprising good earnings. Great growth, etc. I sell some, maybe 15 to 20%. A quarter of my position, and then I trail it. Then I just start trailing it like the first close below this purple line. That’s the 10-day moving average. I sell maybe say a quarter or a third of what I still have left. The shares I have left after the ones I bought. Sold into strength, and then I sell another third or quarter when it hits the first close below the 20-day moving average, which is the yellow line. And then the next level is the red line, which is the 50-day, etc. So that’s how I scale out. So that’s what I use these moving averages for. I wait for them to be my stop pretty much. So I wait until the end of the day and I sell it. Then I also have a level. Let’s just take an example right here. I usually use the previous swing’s lows as my stop. Then I would get out at 79. Even though it turns up intraday and closes above this yellow line, I would have to stop. You know, you can get stopped out because you know. I can, so you know. Sometimes stocks, they go, they can go down, you know, 20%.50% in a day, even. You know mid and large cap names can do that sometimes. So, you know, if it hits my last resort stop, it is my last resort stop. I just sell it. If that was helpful, those are my sell rules, sell into strength and I trail the rest.”

Lone

25,761 views • 11 months ago

RLM is the most import foundation of my Pi Harness (other than Pi of course). It's seeded with late interaction retrieval results (thanks to @lightonai for pylate). The Agent initiates it with query then.. 𝐒𝐞𝐭𝐮𝐩 A python REPL is created and seeded with: 1. Late interaction search to pre-filter. Instead of doing top 3/5/10, it's top hundreds of documents. This is set into a `context` variable. 2. Python functions are loaded in to do more searches if `context` variable isn't enough. And to make llm calls with cheaper models in parallel batches. 𝐈𝐭𝐞𝐫𝐚𝐭𝐢𝐨𝐧 𝐋𝐨𝐨𝐩 From there, an LLM iterates in the REPL based on the query. It's just like exploring in a jupyter notebook. The LLM writes prose (like a markdown cell) and code to be run in the REPL each turn. This allows the LLM to sort, filter, and synthesize information. It can fan out and ask smaller models to summarize, combine, contrast, or do anything else to documents to help it understand the data. After several turns the LLM reponds with the final answer. Either because it found the answer, or hit the budget limit. Context as a Python variable, LLM as the programmer, REPL as the runtime. 𝐖𝐡𝐲 𝐃𝐨𝐞𝐬 𝐓𝐡𝐢𝐬 𝐖𝐨𝐫𝐤 1. Richer Shell. Agents (and subagents) work by intermixing code and prose/thinking. But they use static scripts or bash that run and exit and start over each tool call. That's not ideal for exploration and synthesis of data. For that, state is useful to continue building and exploring the data as you learn more. There's a reason jupyter notebooks have been popular with data scientists. 2. Keeps main agent context clean. The better context you have the better the agent will perform (duh!). This means three thing: better human input, less missing search results, and less incorrect search results. Letting the agent iterate allows it to synthesize just what is needed and nothing else. All bad paths or peeks at something that turns out to be irrelevant stays out of main agent context. 3. Stack the good ideas! People often compare late interaction search vs RLM. Or static vs dynamic languages. Or agentic search vs semantic search. But...You can just use them all together for what they're each good at. Use them all for the area they're really great for. Read the full post which has more detail about how and why.

Isaac Flath

42,620 views • 5 months ago

The same kinds of productivity gains we've seen in coding with AI agents are heading to the rest of knowledge work. This is the jump when you go from having a chatbot to being able to actually have an agent go off and do work for minutes or even hours and come back with a complete work output that you then review. Here's an example of the new Box Agent filling out an RFP response from an existing knowledge base. This process would normally take hours to fill out, and requires the full attention of the user doing the work. Now, you provide the Box Agent with the RFP questions, and it will go off, make a plan, extract all the relevant questions, read through existing source material to come up with an answer, and then generate a new word document as the final output. All while you're doing something else. The key to this architecture is that the agent is able to use all of the same tools in the background that a user uses to get work done. The agent can search for documents, read entire files, run scripts and tools in the background, and even be able to write code on the fly to automate tasks it hasn't seen before. And best of all, the Box Agent will (soon) work from the Box MCP and CLI so you can invoke it in any agentic system as a step in a process. This kind of agent complexity would have been impossible even 6 months ago. Models consistently failed at tracking long running tasks or using the right tools at the right moment for the task. But this is all now possible because of models like GPT-5.4, Opus 4.6, and Gemini 3, and is only getting better by the month. Just as we moved from engineers writing code and using AI as an assistant to answer questions, in many areas of knowledge work -like legal, finance, consulting, sales, marketing, and more- when we have a problem we'll just kick off the AI agent to just go work on it for us in the background.

Aaron Levie

24,728 views • 5 months ago