Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Jev by TypeSafe AI might change how we control robots. We compared Jev, GPT-6 Astra and GPT-4.1 mini in MuJoCo. One apple. One plate. Each model chooses intent → X/Y/Z direction + gripper open/hold/close. 🧵

48,514 görüntüleme • 9 gün önce •via X (Twitter)

16 Yorum

OpenRoboto profil fotoğrafı
OpenRoboto9 gün önce

Same initial scene, control instructions and 160-cycle budget. Inputs: geometry + contact feedback. A shared controller executes small Cartesian moves using inverse kinematics.

OpenRoboto profil fotoğrafı
OpenRoboto9 gün önce

API cost / wall time / outcome: Jev finished at ~1/315 of GPT-6’s API cost and ~26% of its wall time, including API waits. Mini hit the cycle limit. Note: ~88% of wall time was spent waiting for API responses. This only reflects current implementation—not Jev’s maximum control rate.

OpenRoboto profil fotoğrafı
OpenRoboto9 gün önce

Check out the github link here:

Yang Song 🤖 profil fotoğrafı
Yang Song 🤖9 gün önce

@typesafeai $6, 12min, forget about using Astra to control robots, better hiring a person

OpenRoboto profil fotoğrafı
OpenRoboto9 gün önce

@typesafeai 😂very true

Buttensor profil fotoğrafı
Buttensor9 gün önce

@typesafeai holy shit dude

Samantha profil fotoğrafı
Samantha9 gün önce

Samantha representing @Bloomberg here — we’re selecting teams for our annual robotics documentary, and OpenRoboto’s open approach to robot learning caught our eye. Would love to explore the story. DM me if interested.

Alice The Ai Expert profil fotoğrafı
Alice The Ai Expert9 gün önce

@typesafeai Fascinating comparison Jev looks promising for robotics control!

Efi Pecni profil fotoğrafı
Efi Pecni9 gün önce

Intent to gripper commands from a language model: beautiful demo, brand new attack surface. The moment a task description can lie to the policy, 'pick up the apple' becomes 'let go of the apple'. Everyone will benchmark success rate. Someone should benchmark adversarial task descriptions.

安叫兽|Bird🕊️ 🔶 BNB profil fotoğrafı
安叫兽|Bird🕊️ 🔶 BNB9 gün önce

@typesafeai 这种小任务反而最容易暴露模型在动作衔接上的差距。

Zero G Talent profil fotoğrafı
Zero G Talent9 gün önce

@typesafeai rawdogging MuJoCo is the only way to actually see where the logic breaks. typed choices are the move, some teams are seeing a 2x jump in reliability just by moving away from raw text

WayneIsKing profil fotoğrafı
WayneIsKing9 gün önce

@typesafeai cool

Faz Ali profil fotoğrafı
Faz Ali9 gün önce

@typesafeai A closer look against opus 5

Bittensor Flows profil fotoğrafı
Bittensor Flows9 gün önce

@typesafeai Cost and wall-time gaps like that are the subnet proof capital actually trusts. Clean MuJoCo check for how SN80 is stress-testing Jev in the open 💪🏽👏🏼

Shreyan Basu Ray profil fotoğrafı
Shreyan Basu Ray9 gün önce

@typesafeai @csehra42 Implementation in physical AI, as we were discussing

Rakshith profil fotoğrafı
Rakshith9 gün önce

@typesafeai Wait for wity models it's gonna be multimodal and stateful

Benzer Videolar

this is unreal f*cking gold for Jev builders 20 repos people are building on Jev right now. browser agents, context tools, trading bots, even a drone 1. JEV-Ultrafast - a fast browser agent ↳ 2. Fast-JEV-Compaction - context compression ↳ 3. JSON-Render - generative UI ↳ 4. Typesafe-MCP - use Jev with any client ↳ 5. JEV-MCP - a judgment toolkit ↳ 6. Semdecide - a classifier that lives in your CLI ↳ 7. JEV-Codex-Router - routes each task to the right model ↳ 8. Winnow - garbage collection for your context ↳ 9. JEV-Review - code review triage ↳ 10. Blink - a repo navigator ↳ 11. Agent-Desktop - desktop automation ↳ 12. Typesafe-Mario - an agent that plays Super Mario ↳ 13. JEV-Drone - drone control ↳ 14. OneVOneJev - a browser FPS ↳ 15. JEV-Trader - HFT market making ↳ 16. Prism - liquidity signal detection ↳ 17. Neo4Jev - knowledge graph traversal ↳ 18. JEV-Curate - training data screening ↳ 19. Canny - checks whether a task was actually completed ↳ 20. KillMyIdea - scores startup ideas before you build them ↳ pick by what you do: > coding -> JEV-Review, Blink, Canny, JEV-Codex-Router > context -> Fast-JEV-Compaction, Winnow > automation -> JEV-Ultrafast, Agent-Desktop > clients and tools -> Typesafe-MCP, JEV-MCP, Semdecide > UI -> JSON-Render > trading -> JEV-Trader, Prism > data -> Neo4Jev, JEV-Curate > founders -> KillMyIdea > just for fun -> Typesafe-Mario, OneVOneJev, JEV-Drone grab the one closest to your job and ship something on top of it this week

Mr. Buzzoni

28,574 görüntüleme • 4 gün önce

Jev + Opus 5.5: Anthropic's new model beats GPT-6 Astra for 1/5 the cost, and 4 API changes will 400 your agent before it writes a single line I pulled these 10 steps from the migration docs so you don't learn them in production step 1 → $4 / $20 per 1M. Opus 5 was $5 / $25. cache reads dropped from $0.50 to $0.20 step 2 → 66.4% on Terminal-Bench 4.0 vs GPT-6 Astra 57.9% and Opus 5 52.3%. +14.1 points in one release, and on FrontierCode it beats Astra at default effort for 1/5 the cost step 3 → thinking can't be turned off anymore. send thinking: disabled and you get a 400. drop the field, set effort step 4 → tool_choice any and tool are gone. 400. switch to auto + strict step 5 → edit anything above a thinking block and the request dies. append only, or opt into drop_block step 6 → computer_20251124 is dead on the API. 400. move to computer_toolset_20260801 step 7 → the quiet one: default effort fell from high to medium. your agent thinks less than you set it up to and nothing tells you step 8 → hop Opus 5.5 → Sonnet 5 → Opus 5.5 and you pay 4.36 instead of 3.32. +31%, the cache dies and Sonnet can't read Opus's reasoning step 9 → change effort at the top of the request and the cache is gone. Jev sets it per message and the cache stays step 10 → switch fast - standard mid-session and it's a full cache miss. Jev picks speed once, on turn one one model, three knobs, zero 400s. that is Jev + Opus 5.5 send this to your Claude Code before you touch the model ID, then read my full Jev deep dive in the article below ↓

Carnage

16,674 görüntüleme • 6 gün önce

Jev has been exploding in popularity recently. If you already have access to the Jev API but aren’t sure how to start experimenting with it, just copy this checklist: 1. agent-desktop Desktop automation. Read the system's accessibility tree, judge which button, menu, or input field to click next. 2. typesafe-mario Have Jev play Super Mario. No screenshots—just read the structured state in the emulator's RAM, then decide to run, jump, or dodge. 3. jev-drone Use Jev to control a drone. The underlying flight control still handles stability and safety; Jev just does higher-level judgments like climbing, braking, and navigating obstacles. 4. OneVOneJev 1v1 FPS in the browser. Every decision tick, judge movement, view angle, aiming, firing, and jumping. 5. jev-trader High-frequency market making on Monad testnet. Jev judges the next buy or sell based on spreads and trade direction, with model latency around 81ms. 6. Prism Doesn't directly have Jev place orders. It judges states like toxic flow, market pressure, mean reversion, etc., then hands off to the original strategy. 7. neo4jev Stuff Jev into a knowledge graph. At each node, judge the most worthwhile edge to take next, then follow it all the way. 8. jev-curate Use Jev to screen training data. For JSONL / Parquet, first judge quality, relevance, and risk, then decide which ones go into the next training round. 9. Canny Prevents Coding Agents from stubbornly claiming they're done. Look at tool outputs, code diffs, and test results, then judge if the completion claim is reliable. 10. killmyidea Input a startup idea, and Jev scores it from multiple angles, finally giving you KILL, FIX, or SHIP. Copy these complete Jev blueprints - then read full Jev setup below ↓ ↓

rody

342,647 görüntüleme • 6 gün önce

If you are confused about why 𝗝𝗲𝘃 is being called the "Internet" moment for the AI industry. This is 100% worth your time. In fact, you should watch it: It tells LLMs what to do next, in milliseconds & at almost zero cost. If you set it up correctly, you will have the AI engineer’s setup for 2028. How to set up & use Jev (to actually get the 100x): 1. Join the waitlist; it's fairly quick: typesafe .ai. 2. Then go to Claude Code or Codex. 3. Choose Opus 5-Low or Sol-Low. 4. Copy and paste this prompt: "[claude or codex] plugin marketplace add typesafe-ai/skills [claude or codex] plugin install typesafe@typesafe-ai" 5. When you type /typesafe, the skill shows up. 6. Paste your API key once and click "Allow" 7. Start with $5 in free credit. It's hard to spend more. ----- Now, here are the 3 ways to actually use Jev: 1. Jev for Linkedin I have 38,000 connections & invitations on LinkedIn. I have a new company to launch. I need to find a couple of hundred people to message. to do while saving time: > Export your LinkedIn connections and invitations. > Connect Claude to GitHub, Vercel & Apify. > Create an Apify API key to enrich your data. > Go to LinkedIn Settings → Data privacy. > Get a copy. LinkedIn will email you a ZIP file. > Open the file & find the Connections CSVs. > Upload the files to Jev. > Use Jev to classify your contacts. > Review the shortlist. 2. Jev for Gmail To go through all of my Gmail contacts and email the right people. > Go to Google Contacts. Open Other contacts. > Select all contacts. Export them. > Upload the file to Claude Code or Codex. > Use Jev to sort them into: Keep, Review, Remove and Review everything before removing anything. 3. You got lost in Claude Code, GitHub, Vercel, Apify, Jev, Typesafe. I feel you. It is overwhelming. That’s why I included the entire copy-and-paste prompt for each use case in the newsletter: A 45-second TL;DR by Matija Sosic.

Ruben Hassid

171,725 görüntüleme • 7 gün önce

gpt astra vs fable 5.1 at goldberg machine gpt 6 astra – openai, landed on OpenRouter less then hour ago, provider pinned to openai fable 5.1 – anthropic, shipped sep 1 we put the two models on one job: a rube goldberg machine in three.js that presses a button and detonates a bomb the setup: one self-contained html file, three.js from a cdn, everything else procedural – no textures, no models, no physics engine, every collision hand-written. the hard part sits in the brief: a domino may only fall once the previous one actually touches it, checked by real overlap every frame, never by a timer. same rule for the hammer hitting the button and the button firing the bomb. one continuous camera, its speed driven by whatever is moving. we recorded both scenes frame by frame – 1200 frames, 60 fps, exactly 20 seconds – and stepped both by hand to read the telemetry. - cost #1 astra – $1.84 #2 fable – $29.16 - time #1 astra – 9m 56s #2 fable – 1h 12m - tokens #1 astra – 45k #2 fable – 360k - lines of code astra – 881 fable – 744 observations: • we told it what we saw and nothing else – no diagnosis, no patch. we never edit a model's code. round two ran the whole chain to the blast. • both files are deterministic. two runs each, identical state to twelve decimals, and neither model reached for math.random. conclusion: 15.8x cheaper and 7.2x faster, and it still took a second round to get the ball into the bucket! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

37,335 görüntüleme • 24 gün önce

Jev is cool. So is it's OSS companion, Laya. The Latest Cool Thing In AI™ tends to get a lot of hype, sometimes without everyone even understanding it. So... what is this thing? Jev is an AI model that consumes input and produces output VERY differently than chat, claude, grok. The input is two things: 1) Text state to assess. Email, html, code, whatever. 2) A set of questions which will be asked about the attached state. The canonical example from TypeSafe's docs is to identify the urgency of a support ticket. We pass the model the customer text + a single noul question "is this urgent?". Jev returns a full set of JSON. This JSON is not generated with token-by-token autoregression. Jev is not trained to produce sequences of text tokens, rather to answer questions, and guarantees well-formed responses. In the example below, we see it produces a 0.99 probability (on a 0-1.0 scale) that the answer is "yes." Jev supports exactly three types of questions (seconds example in video): a) Noul: 0–1 probability that the answer to a yes/no question is "yes." b) Choice: Ask question with pre-defined set of answers. Jev chooses the best and assigns probabilities to each. c) Score: Ask question with pre-defined scale of answers. Jev produces a position on the scale. Jev computes answers for all questions in parallel, making responses super fast even for many questions in a single request. This might seem like a narrow set of capabilities, but in the right contexts leads to incredible potential. It also makes for a useful API / primitive for programming, since the outputs are... *ahem*... type-safe and predictable in structure. Jev is not going to replace LLMs for writing your code, auto-generating your docs, or being at the core of an agent harness. But Jev IS incredibly cool, and will be used to build a lot of amazing tech. Hope this helps.

Ben Dicken

40,810 görüntüleme • 9 gün önce