Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Jev is actually insane. We benchmarked it, and the results completely change the game for us: > 500 real-time agents > running in parallel > in a 3D environment The preliminary results are crazy: > 500ms average latency > 35 API calls/s > all with a naive implementation We...

58,675 Aufrufe • vor 3 Tagen •via X (Twitter)

21 Kommentare

Profilbild von PSYMOD
PSYMODvor 3 Tagen

Working on real world OSM reasoning and decision making with JEV ! CRAZY OUTCOMES

Profilbild von Cris Lenta
Cris Lentavor 3 Tagen

only a preview whats coming will be insane

Profilbild von PG
PGvor 3 Tagen

Maybe this is the start of an opening into micro-bot biological monitoring / repairs though

Profilbild von Cris Lenta
Cris Lentavor 3 Tagen

tell me more

Profilbild von Dae
Daevor 3 Tagen

Cooooooool

Profilbild von Cris Lenta
Cris Lentavor 3 Tagen

v coooooool

Profilbild von Alexander Goslin
Alexander Goslinvor 3 Tagen

This was one of the first applications that came to mind for Jev. Super cool!

Profilbild von Cris Lenta
Cris Lentavor 3 Tagen

maybe we should work together

Profilbild von PG
PGvor 3 Tagen

Don't start making real-life nanobot swarms 😑

Profilbild von Cris Lenta
Cris Lentavor 3 Tagen

no nano only real-size

Profilbild von Sebastiano Mandalà
Sebastiano Mandalàvor 3 Tagen

They shouldn't release insane AI, so dangerous

Profilbild von Michael Waitze
Michael Waitzevor 3 Tagen

The latency improvements for parallel agents are going to be a total game changer. We actually went deeper on this here:

Profilbild von CuriousDev
CuriousDevvor 3 Tagen

Isn't this close to the state space models, we can remember manipulate and update states this definitely looks interesting

Profilbild von Sacha Martinelle
Sacha Martinellevor 3 Tagen

Impressive! How does that compare to the previous top model?

Profilbild von A Eye Bubble
A Eye Bubblevor 3 Tagen

No you're insane for falling for this scam. Hahahshaha! You're dum

Profilbild von Geωrge
Geωrgevor 3 Tagen

so basically now we have the haves and havenots in low latency token generation

Profilbild von Gregor
Gregorvor 3 Tagen

500ms average is the easy number for multi-agent. genuinely asking what does p99 look like when all 500 agents cluster in the same region at once?

Profilbild von mblaso
mblasovor 3 Tagen

@crislenta what’s insane is that you haven’t posted it on

Profilbild von dp
dpvor 3 Tagen

Probably great for game agents

Profilbild von Justin He
Justin Hevor 3 Tagen

Because JEV is type-safe, doesn't this have limitations on the choices and directions these simulation steer towards? aka not completely free-will for each agent (lol)

Profilbild von AI News Daily
AI News Dailyvor 3 Tagen

Parallel agents are impressive, but the quality checks will decide how useful this is. I would love to see the error rate under real workloads.

Ähnliche Videos

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on Hugging Face, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

kwindla

330,647 Aufrufe • vor 24 Tagen