Loading video...

Video Failed to Load

Go Home

Tetris bench: Jev vs. Laya-mlx I run the Laya model on my MacBook. It's true that the model is really fast (~84ms), but it's also much more stupid. Losing to Jev 3 out of 3 rounds.

74,047 views • 6 days ago •via X (Twitter)

33 Comments

Tony Dinh's profile picture
Tony Dinh6 days ago

Source:

KC's profile picture
KC6 days ago

why only run 3 times, at these costs you should run it at least 100x just to confirm

Tsagaanbayar the Mobile dev's profile picture
Tsagaanbayar the Mobile dev6 days ago

You gotta try mine 👀 Give enough game sense to Laya and it performs great for sure.

Herakles's profile picture
Herakles6 days ago

Jev is just too good

Tobias Wupperfeld's profile picture
Tobias Wupperfeld6 days ago

Yes, I got the same impression. Faster but dumb

Dellort's profile picture
Dellort6 days ago

... that's a fail of comparison. Don't compare timing, compare how many pieces they dropped until the end.

Dmitry's profile picture
Dmitry6 days ago

same here:

Elijah Bare's profile picture
Elijah Bare6 days ago

Similar experience when trying to run on laya - gets stuck. still pretty cool tho!

Sasha Sheng (Hiring) 🫶🏼's profile picture
Sasha Sheng (Hiring) 🫶🏼6 days ago

Aww but hey Jev is punk I mean pink

om's profile picture
om6 days ago

macbook beating the datacenter is my favourite genre

Patrick Daether's profile picture
Patrick Daether6 days ago

Yes, I noticed the same.

Hunter Guo's profile picture
Hunter Guo6 days ago

This is a useful reminder that latency and capability are separate axes. For local agents, the product question is whether the fast model is good enough for routing and control loops, while a stronger model handles the hard turns.

Sulaiman Vesal's profile picture
Sulaiman Vesal6 days ago

84ms per move is useless if the move is wrong. Raw latency wins benchmarks; the interesting question is how far quality scales with a slightly bigger local model.

Kryptoatom ⚡AI Builder's profile picture
Kryptoatom ⚡AI Builder6 days ago

mlx quantizations drop reasoning precision pretty fast once the board gets crowded.

Nicolas Jack's profile picture
Nicolas Jack6 days ago

Noticed Laya gets 6 heuristic-ranked placements in the repo, while Jev sees the full set. Could you try both with the same 6 candidates, plus a heuristic-only baseline? Curious how much of the gap survives with matched inputs.

RΞNΔT⚡️'s profile picture
RΞNΔT⚡️6 days ago

This is exactly what people need to see faster doesn’t always mean better

@bluecow 🐮's profile picture
@bluecow 🐮6 days ago

its kinda clear Laya dont know the game

Julien Active's profile picture
Julien Active6 days ago

84ms but 0-3. Classic speed vs quality tradeoff. My take: the winning setup is routing. I'd run Laya locally for the easy moves and call Jev for the hard ones. Best of both.

Sam's profile picture
Sam6 days ago

Will be interesting to see how much it can be optimised, or trained to specific tasks

Iliyan Topchiev's profile picture
Iliyan Topchiev6 days ago

Speed is ego; depth is the throne. A fast dummy still loses the game.

Maxence's profile picture
Maxence6 days ago

Speed gets attention, but the intelligence gap is the real constraint. Fast wrong moves only tighten the failure loop.

Justin's profile picture
Justin6 days ago

Kev-4B strikes the right balance for me. I agree laya just seems to make too many mistakes

Jatin Garg's profile picture
Jatin Garg6 days ago

84ms is fast, but if Jev solves the Tetris mechanics and Laya doesn't, you're comparing different games. What specific moves does Laya miss?

Carol Rehor's profile picture
Carol Rehor6 days ago

Tetris bench where Laya is fast on the MacBook but Jev still wins the actual game is the eval people needed.

Seed42's profile picture
Seed426 days ago

54ms to stack it wrong faster

Reid Marlow's profile picture
Reid Marlow6 days ago

Sub-hundred millisecond generstion on Apple Silicon makes local quantized weights feel immediate for single-token prediction. Spatial games like Tetris break lightweight models because placing a tetromino requires multi-step lookahead across grid colume heights rather than forward syntax completion. Speed masks the lack of planning depth until the board state demands evaluating branching placements.

Srinivasan K K's profile picture
Srinivasan K K6 days ago

Here we go with another open weight model that competes with Jev seems!!

festus wokoma's profile picture
festus wokoma6 days ago

I guess slow and steady wins the race 🤣

Le Dev ULTIME 🍜's profile picture
Le Dev ULTIME 🍜6 days ago

Just to be sure, did you put that model remotely? Because of you're comparing Jev through API vs a local model... I don't see that as a fair comparison.

josh 🏔️'s profile picture
josh 🏔️6 days ago

more benchmarks like ths

Mikhail Rogov's profile picture
Mikhail Rogov6 days ago

84ms is nothing if it plays worse. speed only matters once accuracy is close, otherwise its just losing fast

onecodeman's profile picture
onecodeman6 days ago

Still needs more training then. Thanks for testing

tteokl's profile picture
tteokl6 days ago

So far this Jev thing is no differenct than a classifier and perform even worse, im confused

Related Videos

Jev is cool. So is it's OSS companion, Laya. The Latest Cool Thing In AI™ tends to get a lot of hype, sometimes without everyone even understanding it. So... what is this thing? Jev is an AI model that consumes input and produces output VERY differently than chat, claude, grok. The input is two things: 1) Text state to assess. Email, html, code, whatever. 2) A set of questions which will be asked about the attached state. The canonical example from TypeSafe's docs is to identify the urgency of a support ticket. We pass the model the customer text + a single noul question "is this urgent?". Jev returns a full set of JSON. This JSON is not generated with token-by-token autoregression. Jev is not trained to produce sequences of text tokens, rather to answer questions, and guarantees well-formed responses. In the example below, we see it produces a 0.99 probability (on a 0-1.0 scale) that the answer is "yes." Jev supports exactly three types of questions (seconds example in video): a) Noul: 0–1 probability that the answer to a yes/no question is "yes." b) Choice: Ask question with pre-defined set of answers. Jev chooses the best and assigns probabilities to each. c) Score: Ask question with pre-defined scale of answers. Jev produces a position on the scale. Jev computes answers for all questions in parallel, making responses super fast even for many questions in a single request. This might seem like a narrow set of capabilities, but in the right contexts leads to incredible potential. It also makes for a useful API / primitive for programming, since the outputs are... *ahem*... type-safe and predictable in structure. Jev is not going to replace LLMs for writing your code, auto-generating your docs, or being at the core of an agent harness. But Jev IS incredibly cool, and will be used to build a lot of amazing tech. Hope this helps.

Ben Dicken

40,810 views • 8 days ago

this is unreal f*cking gold for Jev builders 20 repos people are building on Jev right now. browser agents, context tools, trading bots, even a drone 1. JEV-Ultrafast - a fast browser agent ↳ 2. Fast-JEV-Compaction - context compression ↳ 3. JSON-Render - generative UI ↳ 4. Typesafe-MCP - use Jev with any client ↳ 5. JEV-MCP - a judgment toolkit ↳ 6. Semdecide - a classifier that lives in your CLI ↳ 7. JEV-Codex-Router - routes each task to the right model ↳ 8. Winnow - garbage collection for your context ↳ 9. JEV-Review - code review triage ↳ 10. Blink - a repo navigator ↳ 11. Agent-Desktop - desktop automation ↳ 12. Typesafe-Mario - an agent that plays Super Mario ↳ 13. JEV-Drone - drone control ↳ 14. OneVOneJev - a browser FPS ↳ 15. JEV-Trader - HFT market making ↳ 16. Prism - liquidity signal detection ↳ 17. Neo4Jev - knowledge graph traversal ↳ 18. JEV-Curate - training data screening ↳ 19. Canny - checks whether a task was actually completed ↳ 20. KillMyIdea - scores startup ideas before you build them ↳ pick by what you do: > coding -> JEV-Review, Blink, Canny, JEV-Codex-Router > context -> Fast-JEV-Compaction, Winnow > automation -> JEV-Ultrafast, Agent-Desktop > clients and tools -> Typesafe-MCP, JEV-MCP, Semdecide > UI -> JSON-Render > trading -> JEV-Trader, Prism > data -> Neo4Jev, JEV-Curate > founders -> KillMyIdea > just for fun -> Typesafe-Mario, OneVOneJev, JEV-Drone grab the one closest to your job and ship something on top of it this week

Mr. Buzzoni

28,028 views • 2 days ago

Jev has been blowing up lately. If you've got the Jev API but don't know how to play around with it yet, you can just copy this checklist. 1. jev-ultrafast A high-speed browser Agent built with Browser Use. Jev only judges "what to do, which element to click" at each step, and only calls the small model when typing is needed. Searching for a flight on Google Flights takes about 7 seconds. 2. fast-jev-compaction Context compression for Claude Code. Before each tool call, have Jev judge if there's anything still useful; delete the useless stuff, and keep the original text without rewriting it. 3. json-render Vercel Labs' generative UI framework. In experiments, Jev doesn't write JSON token by token; it just handles selecting components, properties, and layouts. 4. typesafe-mcp Best for people who just got the API. Plug Jev into Claude Code, Claude Desktop, Codex, and Pi, and do Choice / Score / Noul anytime. 5. jev-mcp Ready-made Agent judgment toolkit: fact-checking, content screening, semantic ranking, classification, and information extraction. 6. SemDecide Turn Jev into a command-line tool. Directly classify, score, and filter in the Shell—great for hooking up to crawlers, CI, and data pipelines. 7. jev-codex-router First have Jev judge how hard this round of programming tasks is, then decide the model tier, reasoning depth, and speed mode. 8. Winnow Context garbage collection for Claude Code. When Read / Bash / Grep spits out a ton of stuff, Jev first judges which parts are really relevant to the current task. 9. jev-review Before code review, run it through Jev first to pick out high-risk changes, then hand them off to a pricier big model or a human. Comes with a local dashboard. 10. Blink Use Jev as a code repository navigator. At each directory level, judge which files are most relevant to the current issue, then keep digging down. Copy these complete Jev blueprints - then read full Jev setup below ↓ ↓

rody

199,858 views • 6 days ago