Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

It's impossible for Jev to be good.

101,005 Aufrufe • vor 4 Tagen •via X (Twitter)

37 Kommentare

Profilbild von SpiderMonkey ☀️🌳
SpiderMonkey ☀️🌳vor 4 Tagen

If the job needs speed at low cost jev is the right solution. Start thinking about all the things that need near instant classification. Dont use it where an LLM would work better

Profilbild von kookai · fireply.ai
kookai · fireply.aivor 4 Tagen

what did jev do to deserve this level of confident dismissal lol

Profilbild von quack
quackvor 4 Tagen

obviously its used in combination. it's a specialized model. No shit sherlock it doesn't function like an LLM.

Profilbild von Seth
Sethvor 4 Tagen

I’m sorry, but I 100% would let current models do my taxes and get better results. I already plan on doing that from now on one way or another. The only thing in my way is privacy, so the question is will I let Chinese models do my taxes.

Profilbild von The Human Watch vs AI
The Human Watch vs AIvor 4 Tagen

I too can give a theory of relativity but it will be absolute garbage.

Profilbild von BongBong
BongBongvor 4 Tagen

I've seen a shift in comments so that now "A.I. slop" is just anything someone doesn't like or agree with.

Profilbild von ZazenCodes
ZazenCodesvor 4 Tagen

dude I get the gloom around X demos but Jev is legit. TypeSafe AI has dome something new and exciting here

Profilbild von Mo
Movor 4 Tagen

elaborate

Profilbild von ZazenCodes
ZazenCodesvor 4 Tagen

it pre-trained a neural network to be a general purpose classifier. I do not believe this has been done before also the approach RLCD (reinforcement learning for calibrated decisions) is different than what previous LLMs are doing. aiming to have better alignment with truth rather than "what people want" (e.g. RLFH: chatgpt, claude, etc..)

Profilbild von Mo
Movor 4 Tagen

it’s not uncool. but afaict doesn’t seem even remotely as relevant as llms

Profilbild von J.A. Arroyo
J.A. Arroyovor 4 Tagen

Hey Mo! I agree. To me, Jev has a huge trust problem. At least, with coding, you can run the code and see if it works. With this thing, you get 0.94 and have to blindly trust that it understood the input (context, question, domain, etc.), and that 0.94 actually means something.

Profilbild von René Des States
René Des Statesvor 4 Tagen

"i agree" : 0.9871 "i disagree" : 0.0129

Profilbild von Kent
Kentvor 4 Tagen

Probably good when you can do supplement training on the use case.

Profilbild von Ki
Kivor 4 Tagen

I don't tend to have one way convos, so I prob won't post again. [ I co-created the AI for the X45 to establish my street cred here ] But, I do want to know: how would you measure human intelligence here by these same metrics?

Profilbild von TekJumble
TekJumblevor 4 Tagen

Leads clasification and assignment, cases, spam, product recommendation, rfq routing just to name a few use cases that we can already solve with LLM's but its expensive and its slow this is why this has a market

Profilbild von changdizzlewizzle
changdizzlewizzlevor 4 Tagen

Isn't the issue with your breakdown that it kinda focuses more so on bad implementation than anything else? Like the whole point would be.. let's take CX tickets for example.. you would need to define clear parameters around classification categories for ticket triage

Profilbild von Jeremy
Jeremyvor 4 Tagen

It's a really bad comparison to compare to frontier or even quasi-frontier models. It is amazing at system 1 thinking, and it is exactly a classification-type model. A really good use case is labeling a firehose of data in real time for super cheap, reading a bunch of data and routing it to different places. You can think of it like quasi-reasoning in code. It does not replace any real reasoning steps for AI.

Profilbild von notNaél
notNaélvor 4 Tagen

People like this guy are ignoring the fact that the breakthrough here is the determinism in a type structure and not the intelligence. Intelligence will get better over time. That happens every single time with every model. Be patient.

Profilbild von Ahmed Salem
Ahmed Salemvor 4 Tagen

Sorry pro, but this is a useless video!

Profilbild von Bojan Sala
Bojan Salavor 4 Tagen

That slop thing can be optimized by using a random number generator instead of Jev.

Profilbild von Rohit Pujari
Rohit Pujarivor 4 Tagen

There are no benchmarks for this yet. That’s likely why the quality question is hard to answer. To each their own.

Profilbild von professah X
professah Xvor 4 Tagen

Impossible? Nah. But the amount of training data that it would need to be ACROSS THE BOARD GENERICALLY VIABLE is insane. With blackbox training behind an API, youre not gonna get that. The things you want out of a classifier model - auditing, analysis, consistency - you cant get from Jev. You can, however, quickly classify easy to understand data where a LLM API would be overkill. Soooo... a classifier for vibecoded slop?? 😂😂😂

Profilbild von Nadim
Nadimvor 4 Tagen

I was asking the same question here. I am still waiting for someone to answer:

Profilbild von Panos Daras
Panos Darasvor 4 Tagen

Finally a proper review of this. So much malarkey around this model. (not AI slop comment btw)

Profilbild von lowpass ⚾️
lowpass ⚾️vor 4 Tagen

hang it up broski

Profilbild von Averrouz
Averrouzvor 4 Tagen

It is `Lexical rerranker with dynamic schema apis` I can see some of the use cases but for my agentic work still, I tried many cases and benchmarks not worth my coding or agentic workflows

Profilbild von The Engineer
The Engineervor 4 Tagen

I have been saying the same thing that this is basically classic ML

Profilbild von Ojisan Kaichou
Ojisan Kaichouvor 4 Tagen

Tell us you missed the point without telling us you missed the point.

Profilbild von Shez Malik
Shez Malikvor 4 Tagen

ask it if it's good and it just says true

Profilbild von Siim Haugas
Siim Haugasvor 4 Tagen

using it for trading is especially ridiculous because it's a text classifier. it's genuinely good at that: 99% on exchange notices when it's confident. ask it where price goes and it's worse than a a coin flip: 48.4% over 298k trades. even TypeSafe doesn't claim it predicts prices. source: backtested

Profilbild von webXOS
webXOSvor 4 Tagen

"good" being vague in it's own right: Jev is a tool more than a model tbh - the marketing team deserves a raise lol

Profilbild von jaysouthbets
jaysouthbetsvor 4 Tagen

you follow Mo @Abomination81 ? I like his various insights.

Profilbild von Sal Iozzia
Sal Iozziavor 4 Tagen

the classification quality is also up to your implementati9on and what you are defining as its base of data to make the comparison from.

Profilbild von Robert Sale
Robert Salevor 4 Tagen

Yo, all I gotta say is I desperately need that slop stamp thing 😂

Profilbild von curiousNerd
curiousNerdvor 4 Tagen

Finally someone mentioned about “machine learning” ! Thank you 🙏 you will be remembered ..

Profilbild von Sooraj Chandran
Sooraj Chandranvor 4 Tagen

It might be because we are evaluating it against wrong use case? Agree it shouldn't be used for "judgement" - but for a lot of simpler classification, it's a great solution. Most companies can use it without having to worry about fine-tuning or hosting their own models.

Profilbild von 1Broom
1Broomvor 4 Tagen

Impossible, sure. It's just been catching stuff in my long test runs that the green checkmarks swore wasn't there. Totally not good though.

Ähnliche Videos