Загрузка видео...

Не удалось загрузить видео

На главную

VERCEL + CLOUDFLARE + CLAUDE CODE all picked up Jev in 6 days... an AI that can't write a single word became the fastest adopted model on Vercel's AI Gateway Jev decides → big model thinks → code authorizes TypeSafe dropped it on September 15, founder Diogo Almeida worked...

16,184 просмотров • 7 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

this is pure f*cking treasure. my ai bill last month: claude max 20x ............ $200 chatgpt pro ............... $200 supergrok heavy ........... $300 total ..................... $700 then i went through github and found 5 repos that eat the boring half of that bill on my own machine rizzo-flow ................ 833 stars the open local take on typesafe's jev. a 4.4gb model on llama.cpp answers yes / no / pick one / score with a probability on every answer, and you point your api at localhost by changing one url llm2jev ................... 391 stars turns qwen3.5-4b into a jev-style decision model. 48ms p50 in their own benchmark, against 652ms for the real jev fast-browser-use .......... 202 stars a browser agent that runs 100% local on qwen3.5-9b and plugs into claude code and codex as a skill. opens the right wikipedia article in about 4 seconds, zero cloud calls deepseekgui ............... 84 stars a desktop workbench on deepseek harness with git, a built-in browser and memory. you pay deepseek per token instead of a flat $200 a month oriveo .................... 25 stars one app for openai, anthropic, gemini, grok, deepseek and 10 more providers, plus ollama on your own gpu. your keys, no subscription, no account claude still writes my hardest code. the yes or no calls, the browser clicking and the everyday chatting moved to my gpu and my own api keys, and the receipt stopped looking like rent repos in the comments

starmex

73,538 просмотров • 3 дней назад

Jev dropped and people immediately started doing stupid sh*t with it people gave this thing trading bots, browsers, Doom, drones, dinosaurs and tax forms and basically said: “you decide.” here’s what happened: > jev-trader Jev gets a new Monad block every ~300ms and decides whether to place a live limit order. no essay. just the decision. 1,911 stars ▸ ⁠ > jev-ultrafast a browser agent where Jev decides every click. the text model only wakes up when something actually needs to be written. 16,758 stars ▸ ⁠ > jev-doom-agent someone compiled actual Chocolate Doom to WebAssembly and let Jev make the tactical decisions every frame. yes. Doom. ▸ ⁠ > jev-t-rex-runner remember that stupid Chrome dinosaur? now Jev plays it. jump → duck → run → repeat ▸ ⁠ > typesafe-chess Jev went against a real chess search engine. two games. colors swapped. the search engine won both — and overruled Jev’s first instinct on roughly half the moves. ▸ ⁠ > jev-drone camera → Jev → drone. a simulated quadrotor has to clear five stations while Jev looks at the situation twice a second and decides what happens next. ▸ ⁠ > tax-doc-classifier then someone gave it IRS forms. 261 documents. 100% strict accuracy. roughly $0.001/page. ▸ ⁠ > killmyidea this one is evil lol tell it your startup idea. Jev looks at it from different angles and gives you: KILL / FIX / SHIP ▸ ⁠ > jev-curate throw huge Parquet / JSONL datasets at it. Jev judges 1,500+ rows/sec and keeps only the stuff that passes your rules. ▸ ⁠ > pg-jev and now it’s inside PostgreSQL. ask questions about your own tables in plain English → get the decision back. ▸ ⁠ and this is the weird part: none of these need Jev to write you a beautiful paragraph. they need it to look at a situation and pick: BUY / WAIT CLICK A / CLICK B JUMP / DUCK KILL / FIX / SHIP KEEP / DROP that’s basically the whole Jev idea. LLMs think and write. Jev decides. code does. full setup + my three-question Jev test below

kiosa

67,168 просмотров • 15 дней назад

HydraFusion Explained. Part I: How does the Copilot engine know what to optimize for? Your prompt is evaluated across 4 dimensions: ➡ Does it require deep reasoning? (aka. reasoning depth) ➡ Is it a sophisticated problem? (aka. code generation complexity) ➡ Is it untangling a complicated mess? (aka. debugging difficulty) ➡ Is it dominated by tool-use? (aka. tool orchestration needs) Based on this evaluation, a HyDRA score is assigned to determine the capability profile your task needs the most and to establish a quality bar. Part II: How does it choose a model? Note: It doesn' t pick one model to handle the entire job e2e, (that's Auto mode). Instead, it selects 1 of 3 execution workflows and assigns the best model at different stages based on the HyDRA score: 1️⃣ Single ⚙️ How it works: A single model completes the task from start to finish. ⚖️ Rationale: The task comfortably meets the quality bar with one model. Multi-model orchestration would add latency and cost with no meaningful quality gain. 2️⃣ Cascade ⚙️ How it works: A lightweight, cost-efficient model generates the solution. This draft is evaluated against a quality gate and if it falls short of the quality bar, the entire task escalates to a stronger, frontier model. ⚖️ Rationale: Only bring in the big guns when there is concrete evidence that a lightweight model won't meet the quality threshold. 3️⃣ Critique ⚙️ How it works: A lightweight model drafts the initial code and tool interactions. An independent, read-only frontier model reviews that draft and provides feedback. The original lightweight model then performs any targeted revision(s) before the final response is sent to the user. ⚖️ Rationale: Writing code (output tokens) is expensive while reviewing code (input tokens) is cheap. Instead of incurring the cost of a powerhouse writing hundreds of lines from scratch, a cost-efficient model writes the first draft, and the frontier model just reviews it and points out fixes. HydraFusion is available in experimental preview on the GitHub Copilot CLI: /experimental on, /model and select Hydrafusion (Research Preview)

Julia Muiruri

13,000 просмотров • 1 месяц назад

Jev + SERV is actually insane. We already showed you can increase Jev's performance with SERV Reasoning. Now we're taking it further, bringing Jev-powered Decision nodes into Graph Sharding with the upcoming SERV v3. Here's a breakdown of how it works: Jev is a decision-making model. Given a task and a set of options, it predicts which path is more likely. Think of the octopus that predicted World Cup results. Jev does that for your business, except it's not luck. It weighs every option and tells you how sure it is. It does this by assigning probabilities to outcomes. It doesn't generate text on its own, so you can't expect it to create a new outcome for you. But that's also what enables it to be lightning fast and dirt cheap. For example, in customer service you can ask Jev how to triage an incoming query and route it to the correct department. It can only select from the list of departments you provide it. This also means it can't hallucinate a new outcome outside the options it's given, which makes it incredibly interesting for OpenServ. In Graph Sharding, we take a single system prompt and break it down into multiple LLM steps with deterministic input and output shapes. Some of these steps require an LLM to produce new output, while others are simply decision routers that determine the next possible path. Traditionally, LLMs are slow and expensive. Breaking a single prompt into multiple steps increases accuracy and reliability by a ton, but it also introduces latency. Jev takes on those decision nodes, which are the backbone of a business process and therefore SERV graphs, and makes them super consistent and lightning fast, lowering the overall cost and latency of graph execution. SERV Reasoning on its own is a great force multiplier for Jev because, like all other models, it works by interpreting input instructions. The clearer those instructions are, the better the model performs. That's where SERV Reasoning comes into play. Just like amplifying any other model, we also amplify the accuracy and consistency of Jev's responses. And now we're bringing Jev-powered Decision nodes into Graph Sharding with SERV v3.

Armagan Amcalar

365,108 просмотров • 16 дней назад

I built HypeMeter in 4 hours with Jev + Minds. Its best trick is saying no, and deciding what is likely a rug vs real hype. I am giving away an Argonaut NFT to reward Beta testers. Yes, that's you. Every "alpha bot" screams BUY. None of them tell you which cheap listings are cheap for a reason. So I wired two things together: Jev by TypeSafe AI . It does not write essays. It answers typed questions: pick one, score this, yes or no. About a third of a second per decision, cheap enough to judge every cheap listing instead of a shortlist. Minds by Minds by Animoca Brands . Your own AI agent. Tell it your strategy in plain words ("Argonauts under 0.3, grade A or better") and it messages you one digest a day, pings you whenever steals are available. First full sweep: 898 listings across 20 collections, including Robinhood (of course). Calls that survived: one. And that one was my own bug: an "83% edge" that was a 2-item bid read as one. The sanity check now kills those before anyone sees them. That is the product. Most cheap NFTs are traps, and it says so. It also hunts rares priced under what their trait actually sells for. Yesterday it flagged an Argonaut with a 1-in-70 palette, listed at 0.79 ETH two days before the same palette sold for 0.9 and 1.0. No hindsight. Every call is written down the moment it is made, then graded at 24 hours and 7 days. Public scoreboard, losses included. Free while in beta. Sign in with Minds: And yes, the giveaway is real: Argonaut #2764 goes to someone who actually uses it. Every active day is an entry, there is a leaderboard, and signing in before 24 Sept gets you 3 bonus entries. Rules on the site. RT and comment "Jev" for extra entry. Have fun sniping.

Jesus is Lord | Chev

51,733 просмотров • 17 дней назад

New open-source agent harness just landed! I got early access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.

elvis

11,303 просмотров • 1 месяц назад

I GAVE JEV A CRAWLER + ONE GPT AGENT AND ASKED WHERE THE ONLINE MONEY IS RIGHT NOW it read 5,137 open jobs and found 3 niches where clients pay and almost nobody bids everyone asks chatgpt for side hustle ideas everyone gets the same 10 answers so i made it read what clients are actually paying for this week a crawler is a bot that reads the internet for you jev is the ai that answers every question with a number, under a second, under a cent gpt is the expensive brain, it only gets called when jev says a niche is worth it the crawler reads, jev decides, gpt researches honestly gpt alone just guesses, jev is what turns 5,137 jobs into 3 answers what the three of them did: -> crawler read all 5,137 open jobs on freelancer, budgets, skills, how many people already bid -> jev checked every one: real job? doable online by one person? 684 got cut -> gpt named 30 niches from 600 of them, jev sorted all 4,453 into those niches -> jev looked at each niche's numbers and sent only 12 to gpt -> gpt researched those 12 on the open web: other platforms, real prices, how crowded it is -> 3 came back with real demand and fewer bids than the typical job (14): 01 chrome extension · 85 jobs this week · median $545 · 11 bids each · ~$4,671/mo 02 tiktok shop setup · 91 jobs this week · median $232 · 6.5 bids each · ~$1,984/mo 03 notion setup · 95 jobs this week · median $181 · 3 bids each · ~$1,551/mo potential = your fair share of this week's jobs (budget ÷ (bids + 1)), max 2 jobs a week not every pick held up, excel dashboards looked great on freelancer but gpt found it crowded everywhere else jev made 9,617 calls for $0.13, gpt on every job would have cost ~$58.66, the whole run cost $6.53 i wrote up why the cheap brain decides and the expensive one only gets called when it's worth it, it's below ↓ costs nothing bookmark this, empty niches don't stay empty show this to the one friend who keeps asking chatgpt for side hustle ideas should i run it on upwork next, or is chrome extension the one?

Paone

49,272 просмотров • 4 дней назад

Don't train the model, evolve the harness. I read a brilliant blog post from Hugging Face where they took a frozen open model scoring 0% on a hard legal agent benchmark, left its weights alone, and let an automated loop rewrite only the code around it. That code layer is the harness, the runtime wrapper that feeds the model context, runs its tool calls, and decides when a run ends. By the time the loop finished, the system had essentially matched Sonnet 4.6 on the benchmark's headline metric, at roughly 7x lower cost per task. Zero weights changed. The gain existed because of where the model was failing. The judge only grades files saved in the right place under the exact requested filename, and the model kept doing the legal analysis correctly, then saving it under the wrong name, dropping it in a scratch folder, or never writing it at all. So the 0% was never measuring legal reasoning. It was measuring the harness. Hand-tuning that layer is slow and model-specific, so they automated it. A Claude proposer adds exactly one mechanism per iteration, and an outer loop keeps it only if it clearly beats the current best, so accepted mechanisms compound. What the loop discovered says a lot about where agents actually fail. → The biggest single gain was file handling, not intelligence. An automatic step that lands the deliverable exactly where the judge expects it beat every prompt change, with zero extra model tokens. → Code fixes transferred across models, prompt playbooks did not. The same harness lifted a smaller model from the same family by 14 points, but the tuned prompts hurt a different model family on tasks it could already finish. → The harness mattered more than anything else. Same model, same judge, same tasks, and five different harnesses scored anywhere between 3.5% and 80.1%. The gains do eventually flatten, and the remaining misses look like real capability gaps. At some point the wrapper runs out of tricks and the model has to carry the work. But the lesson holds. A benchmark score measures the model and its harness together, and until the harness is fixed, it's impossible to know which one failed. I highly recommend reading this: I also wrote a deep dive on agent harness engineering a while back, covering the orchestration loop, tools, memory, context management, and everything that turns a stateless LLM into a capable agent. The article is quoted below.

Akshay 🚀

245,379 просмотров • 3 месяцев назад

Claude Code tip: if Opus 5.5 is already your main model, Fable 5.1 has been sitting idle this whole time. wire it in with /advisor start it with /advisor fable Opus 5.5 writes every line. Fable 5.1 reads the whole session, every tool call, and says nothing until one of three moments: → a plan gets proposed: is this actually the right move, or just the first one? → the same error comes back twice: is the search stuck, or is this a dead end? → the task gets marked done: what got missed while it was moving fast? Opus 5.5 ships. Fable 5.1 catches what would've shipped broken. Jev engineering makes the same move one layer down: forks that don't need a real thinker, which file, which tool, retry or give up, get routed to Jev and answered in under half a second. the expensive model only ever sees the forks that genuinely split. the tree this runs on: > Opus 5.5, high effort, owns the main session > explorer, medium effort, reads the code > worker, medium effort, edits and runs tests > researcher, medium effort, pulls the docs > Fable 5.1 outside all of it, on call, never writing a line itself drop the tree and this prompt into Claude Code: "Rebuild my Claude Code setup around this tree: 1. Look in ~/.claude/agents and .claude/agents for subagents that already cover explorer, worker and researcher. Draft new ones only for roles that are missing. Set each to model: opus, effort: medium. If an existing subagent is pinned to a different model, list it, don't touch it. 2. Set the main session's effortLevel to high in ~/.claude/settings.json, and set advisorModel to fable. 3. Check for anything disabling the advisor: CLAUDE_CODE_DISABLE_ADVISOR_TOOL, DISABLE_TELEMETRY, any variable blocking feature-flag fetches, and CLAUDE_CODE_EFFORT_LEVEL, which overrides subagent effort. Report what you find. Change nothing yet. 4. Add one line to ~/.claude/CLAUDE.md: consult the advisor before a large plan, when an error repeats, and before marking a long task done. Show every change as a diff first. Don't touch anything until I say go."

Ryven

35,757 просмотров • 8 дней назад