Loading video...
Video Failed to Load
Deploying DiffusionGemma-Jev (djev) just got a lot easier. You can now spin up a Jev API-compatible endpoint on Google Cloud Run using a single command. Performance is solid: ~35-60 ms for single step latency and batch@32 is ~100-123 requests/sec. It's a straightforward way to experiment without needing your own... show more
627,376 views • 2 days ago •via X (Twitter)
31 Comments

Credits to @mmastrac and @dylayed. Original post:

@bodonoghue85 Fastest bookmark of my life

jev on blackwell in the big 2026 🥀 what are we doing

but why ? jev is cheaper than $3/h and way better at it

Infinite Jev unlocked

Put some pressure on Gemini team so you guys can distill for Gemma 5 soon 😌

this is exactly, what i asked for. thanks!

This is the kind of deployment UX that makes experimentation actually accessible. Spin it up when you need it, pay while it runs, and let it disappear when you’re done. 🔥

When I was seeing Jev explanation I was asking myself, isn't what DiffusionGemma already doing. Is it the similar, is it that Jev just has more hype?

Coming to the Gemini API/AI Studio?

lmao, didnt see this coming

Single-ste latency is the real flex. 35ms means you can feel the stream before your coffee cools, smooth workflow vibes that actually respect a mom’s attention span. Efficient code looks good on everyone. 😊

One practical limit in the repo: the Cloud Run recipe is set to one GPU instance and 32 concurrent requests. The 100–123 req/s figure is a warm batch result; it says little about the wait a new caller sees when that GPU is busy.

google's like you want ML, we got ML

Oh ok Google.

Interesting! How does this compare to running similar models on other serverless platforms like AWS Lambda or Azure Functions in terms of cost and cold start times?

@mmastrac you mad lad

Don't add --cpu-boost by reflex here. The repo reports 73-75s cold starts with it versus 47.5s without it on the 20-vCPU GPU setup.

一条命令就能把实验跑起来 这种少折腾的升级最有用 普通人不用先学一堆黑话

The interesting part is the deployment contract: Jev-compatible API, scale-to-zero, and a single Cloud Run command. The next useful artifact is a cold-start and tail-latency trace under real bursty traffic, because “$0 when idle” matters only when wake-up cost stays predictable.

@grok 这个有什么业务意义呢?

Gemini 4 pro gelene kadar tüm yan görevleri yapıyorlar

OMG GEMMA I LOVE YOU

Single-command deployment is the kind of boring improvement engineers quietly celebrate.

how smart is this though?

Local?

AI is getting so advance its going backwards to 2017. 😂

Heh heh

Need to check if laya is better or djev? 🤔

Is there a blueprint like this for an openai compatible version? Can already get structured output that way. Jev format just seems like nerfing Gemma

google made deploying its diffusiongemma image model a single cloud run command, $3 an hour and zero when idle. the gpu shortage is just a pricing strategy now.

