Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Deploying DiffusionGemma-Jev (djev) just got a lot easier. You can now spin up a Jev API-compatible endpoint on Google Cloud Run using a single command. Performance is solid: ~35-60 ms for single step latency and batch@32 is ~100-123 requests/sec. It's a straightforward way to experiment without needing your own...

627,376 Aufrufe • vor 2 Tagen •via X (Twitter)

31 Kommentare

Profilbild von Google Gemma
Google Gemmavor 2 Tagen

Credits to @mmastrac and @dylayed. Original post:

Profilbild von Jonathan Ouyang
Jonathan Ouyangvor 2 Tagen

@bodonoghue85 Fastest bookmark of my life

Profilbild von Viren Khandal
Viren Khandalvor 2 Tagen

jev on blackwell in the big 2026 🥀 what are we doing

Profilbild von ORC KILLER
ORC KILLERvor 2 Tagen

but why ? jev is cheaper than $3/h and way better at it

Profilbild von Everlier
Everliervor 2 Tagen

Infinite Jev unlocked

Profilbild von hey
heyvor 2 Tagen

Put some pressure on Gemini team so you guys can distill for Gemma 5 soon 😌

Profilbild von iMuffin
iMuffinvor 2 Tagen

this is exactly, what i asked for. thanks!

Profilbild von Vincent Myshinkye
Vincent Myshinkyevor 2 Tagen

This is the kind of deployment UX that makes experimentation actually accessible. Spin it up when you need it, pay while it runs, and let it disappear when you’re done. 🔥

Profilbild von Ibrahima Sarr
Ibrahima Sarrvor 2 Tagen

When I was seeing Jev explanation I was asking myself, isn't what DiffusionGemma already doing. Is it the similar, is it that Jev just has more hype?

Profilbild von Red Wizard 🔮 (e/acc)
Red Wizard 🔮 (e/acc)vor 2 Tagen

Coming to the Gemini API/AI Studio?

Profilbild von Sadnan
Sadnanvor 2 Tagen

lmao, didnt see this coming

Profilbild von The AI Therapist
The AI Therapistvor 2 Tagen

Single-ste latency is the real flex. 35ms means you can feel the stream before your coffee cools, smooth workflow vibes that actually respect a mom’s attention span. Efficient code looks good on everyone. 😊

Profilbild von TensorQuay
TensorQuayvor 2 Tagen

One practical limit in the repo: the Cloud Run recipe is set to one GPU instance and 32 concurrent requests. The 100–123 req/s figure is a warm batch result; it says little about the wait a new caller sees when that GPU is busy.

Profilbild von pysolin
pysolinvor 2 Tagen

google's like you want ML, we got ML

Profilbild von 𓂀♎︎𓂀
𓂀♎︎𓂀vor 2 Tagen

Oh ok Google.

Profilbild von Bruno B. Sant' Anna
Bruno B. Sant' Annavor 2 Tagen

Interesting! How does this compare to running similar models on other serverless platforms like AWS Lambda or Azure Functions in terms of cost and cold start times?

Profilbild von Kyle Kelley
Kyle Kelleyvor 2 Tagen

@mmastrac you mad lad

Profilbild von John Rood
John Roodvor 2 Tagen

Don't add --cpu-boost by reflex here. The repo reports 73-75s cold starts with it versus 47.5s without it on the 20-vCPU GPU setup.

Profilbild von Jhon Dennis
Jhon Dennisvor 2 Tagen

一条命令就能把实验跑起来 这种少折腾的升级最有用 普通人不用先学一堆黑话

Profilbild von 刘朝 Zhao Liu
刘朝 Zhao Liuvor 2 Tagen

The interesting part is the deployment contract: Jev-compatible API, scale-to-zero, and a single Cloud Run command. The next useful artifact is a cold-start and tail-latency trace under real bursty traffic, because “$0 when idle” matters only when wake-up cost stays predictable.

Profilbild von Y11
Y11vor 2 Tagen

@grok 这个有什么业务意义呢?

Profilbild von Karios
Kariosvor 2 Tagen

Gemini 4 pro gelene kadar tüm yan görevleri yapıyorlar

Profilbild von gab
gabvor 2 Tagen

OMG GEMMA I LOVE YOU

Profilbild von Fajar M Reza
Fajar M Rezavor 2 Tagen

Single-command deployment is the kind of boring improvement engineers quietly celebrate.

Profilbild von Hidde
Hiddevor 2 Tagen

how smart is this though?

Profilbild von extracsw
extracswvor 2 Tagen

Local?

Profilbild von Noble Savage
Noble Savagevor 2 Tagen

AI is getting so advance its going backwards to 2017. 😂

Profilbild von 🇺🇸🏴‍☠️
🇺🇸🏴‍☠️vor 2 Tagen

Heh heh

Profilbild von Jigs
Jigsvor 2 Tagen

Need to check if laya is better or djev? 🤔

Profilbild von Chase McDougall
Chase McDougallvor 2 Tagen

Is there a blueprint like this for an openai compatible version? Can already get structured output that way. Jev format just seems like nerfing Gemma

Profilbild von Sonath
Sonathvor 2 Tagen

google made deploying its diffusiongemma image model a single cloud run command, $3 an hour and zero when idle. the gpu shortage is just a pricing strategy now.

Ähnliche Videos

Jev has been blowing up lately. If you've got the Jev API but don't know how to play around with it yet, you can just copy this checklist. 1. jev-ultrafast A high-speed browser Agent built with Browser Use. Jev only judges "what to do, which element to click" at each step, and only calls the small model when typing is needed. Searching for a flight on Google Flights takes about 7 seconds. 2. fast-jev-compaction Context compression for Claude Code. Before each tool call, have Jev judge if there's anything still useful; delete the useless stuff, and keep the original text without rewriting it. 3. json-render Vercel Labs' generative UI framework. In experiments, Jev doesn't write JSON token by token; it just handles selecting components, properties, and layouts. 4. typesafe-mcp Best for people who just got the API. Plug Jev into Claude Code, Claude Desktop, Codex, and Pi, and do Choice / Score / Noul anytime. 5. jev-mcp Ready-made Agent judgment toolkit: fact-checking, content screening, semantic ranking, classification, and information extraction. 6. SemDecide Turn Jev into a command-line tool. Directly classify, score, and filter in the Shell—great for hooking up to crawlers, CI, and data pipelines. 7. jev-codex-router First have Jev judge how hard this round of programming tasks is, then decide the model tier, reasoning depth, and speed mode. 8. Winnow Context garbage collection for Claude Code. When Read / Bash / Grep spits out a ton of stuff, Jev first judges which parts are really relevant to the current task. 9. jev-review Before code review, run it through Jev first to pick out high-risk changes, then hand them off to a pricier big model or a human. Comes with a local dashboard. 10. Blink Use Jev as a code repository navigator. At each directory level, judge which files are most relevant to the current issue, then keep digging down. Copy these complete Jev blueprints - then read full Jev setup below ↓ ↓

rody

194,422 Aufrufe • vor 4 Tagen

Jev is cool. So is it's OSS companion, Laya. The Latest Cool Thing In AI™ tends to get a lot of hype, sometimes without everyone even understanding it. So... what is this thing? Jev is an AI model that consumes input and produces output VERY differently than chat, claude, grok. The input is two things: 1) Text state to assess. Email, html, code, whatever. 2) A set of questions which will be asked about the attached state. The canonical example from TypeSafe's docs is to identify the urgency of a support ticket. We pass the model the customer text + a single noul question "is this urgent?". Jev returns a full set of JSON. This JSON is not generated with token-by-token autoregression. Jev is not trained to produce sequences of text tokens, rather to answer questions, and guarantees well-formed responses. In the example below, we see it produces a 0.99 probability (on a 0-1.0 scale) that the answer is "yes." Jev supports exactly three types of questions (seconds example in video): a) Noul: 0–1 probability that the answer to a yes/no question is "yes." b) Choice: Ask question with pre-defined set of answers. Jev chooses the best and assigns probabilities to each. c) Score: Ask question with pre-defined scale of answers. Jev produces a position on the scale. Jev computes answers for all questions in parallel, making responses super fast even for many questions in a single request. This might seem like a narrow set of capabilities, but in the right contexts leads to incredible potential. It also makes for a useful API / primitive for programming, since the outputs are... *ahem*... type-safe and predictable in structure. Jev is not going to replace LLMs for writing your code, auto-generating your docs, or being at the core of an agent harness. But Jev IS incredibly cool, and will be used to build a lot of amazing tech. Hope this helps.

Ben Dicken

40,810 Aufrufe • vor 5 Tagen

Jev builds the MOST POWERFUL trading agents and someone JUST open sourced jev-trader, a fully working 24/7 trading bot with Jev along with COMPLETE low latency CODEBASE WHAT THIS MEANS FOR YOU - you no longer have to build a trading bot with Jev from scratch, you just clone this and make it yours here is how you make your own Jev trading bot with this repo: 1. clone it and run three commands, it boots straight into dry run mode with real book data, real decisions, and simulated fills so you can watch it think with zero capital 2. drop in your Jev API key and the model starts answering buy or sell on every block with calibrated probabilities in 81 milliseconds 3. swap the book reader for your own venue, the model interface is clean so any order book that returns bids and asks plugs straight in 4. tune the decision cadence and horizon, ask the model every N blocks about the move over the next M, so you control how aggressive the engine trades 5. the hot loop already fits one block with exactly two round trips, one to read the book, one to send the order, nothing else on the path, this is the institutional latency discipline most retail bots never reach 6. plug in the live server and every block, every decision, every fill streams to a public dashboard so you watch your engine run the whole point is this repo hands you HARDEST part for FREE - > the low latency engine the COMPLETE breakdown of how i turned this into hedge fund grade HFT trading system is in my article below:

Roan

119,605 Aufrufe • vor 5 Tagen

If you are confused about why 𝗝𝗲𝘃 is being called the "Internet" moment for the AI industry. This is 100% worth your time. In fact, you should watch it: It tells LLMs what to do next, in milliseconds & at almost zero cost. If you set it up correctly, you will have the AI engineer’s setup for 2028. How to set up & use Jev (to actually get the 100x): 1. Join the waitlist; it's fairly quick: typesafe .ai. 2. Then go to Claude Code or Codex. 3. Choose Opus 5-Low or Sol-Low. 4. Copy and paste this prompt: "[claude or codex] plugin marketplace add typesafe-ai/skills [claude or codex] plugin install typesafe@typesafe-ai" 5. When you type /typesafe, the skill shows up. 6. Paste your API key once and click "Allow" 7. Start with $5 in free credit. It's hard to spend more. ----- Now, here are the 3 ways to actually use Jev: 1. Jev for Linkedin I have 38,000 connections & invitations on LinkedIn. I have a new company to launch. I need to find a couple of hundred people to message. to do while saving time: > Export your LinkedIn connections and invitations. > Connect Claude to GitHub, Vercel & Apify. > Create an Apify API key to enrich your data. > Go to LinkedIn Settings → Data privacy. > Get a copy. LinkedIn will email you a ZIP file. > Open the file & find the Connections CSVs. > Upload the files to Jev. > Use Jev to classify your contacts. > Review the shortlist. 2. Jev for Gmail To go through all of my Gmail contacts and email the right people. > Go to Google Contacts. Open Other contacts. > Select all contacts. Export them. > Upload the file to Claude Code or Codex. > Use Jev to sort them into: Keep, Review, Remove and Review everything before removing anything. 3. You got lost in Claude Code, GitHub, Vercel, Apify, Jev, Typesafe. I feel you. It is overwhelming. That’s why I included the entire copy-and-paste prompt for each use case in the newsletter: A 45-second TL;DR by Matija Sosic.

Ruben Hassid

167,306 Aufrufe • vor 4 Tagen