正在加载视频...

视频加载失败

Look ma, my AI-powered Jax beats Jev and Laya when benchmarked correctly.

15,618 次观看 • 8 天前 •via X (Twitter)

22 条评论

Nandakishor m 的头像
Nandakishor m7 天前

Creator of laya here. Super cool

Ivan Khokhlov 的头像
Ivan Khokhlov7 天前

What kind of benchmark setup are you runing here? Never heard of Jax, is it also focused purely on classification workloads?

Tom Siwik 的头像
Tom Siwik7 天前

Jax is my own experimental classifier - Jev api compatible so I can run it locally. Currently I'm benchmarking differently in Python. This is plain frontend Javascript workfloads benchmarking a game - meaningless but it's been shown around on X, pretending Laya is beating Jev

Ivan Khokhlov 的头像
Ivan Khokhlov7 天前

Thank you Tom! Would love to know more once you make more progress. I am looking now in general for reliable classifiers.

Tom Siwik 的头像
Tom Siwik7 天前

reliable is a loose term with llm-based classifiers. but will let you know

Sasha Sheng (Hiring) 🫶🏼 的头像
Sasha Sheng (Hiring) 🫶🏼7 天前

Very cool! Thanks for using pink for Jev!

Tom Siwik 的头像
Tom Siwik7 天前

It's your color - I fixed it (it's reversed in the video) lol

Saïd Aitmbarek 的头像
Saïd Aitmbarek7 天前

Dope experiment mate. What's the algo/concept behind Jax?

Tom Siwik 的头像
Tom Siwik7 天前

Train a small model to classify. It reads input, but generates nothing (no inferrence). Shows raw numbers what it would have predicted for the schema/question/options

David R. Prasser 的头像
David R. Prasser7 天前

congrats!

Jeremy 的头像
Jeremy7 天前

Is there a known strategy in Snake that suggests you go back to the left wall every couple of seconds? It seems like all 3 models behave that way, even when it was super inefficient (in the early game). Curious what context led to that.

Tom Siwik 的头像
Tom Siwik7 天前

I couldn't tell you why. looks like the training data suggest "left" as the safest bet

Nour Eddine Hamaidi 的头像
Nour Eddine Hamaidi7 天前

ok but where's JAX?

Tom Siwik 的头像
Tom Siwik7 天前

Jax is tugged in and sleeping (still in development). Will be used for code classification and don't think it'll be useful for the general public. (Un)surprisingly albeit trained on code it can still play games

Paul Yorke 的头像
Paul Yorke7 天前

Look ma energy is strong. I’ll take a live walk on my own app over a chart though. Same repo, same flow. That’s the only benchmark I actually trust.

Tom Siwik 的头像
Tom Siwik7 天前

That's just a frontend X-friendly benchmark. I can't look inside Jev tho :( But it's fun to reconstruct it.

Paul Yorke 的头像
Paul Yorke7 天前

Fair. Chart’s for the timeline. The reconstruction’s the actual job.

ÆL|Ξ 的头像
ÆL|Ξ7 天前

Can Jev really plan ahead of the moves for the end game? I really wanted to see the final part for each model? I don't think there is an intelligence in the jev. It is just an classifier model.

Tom Siwik 的头像
Tom Siwik7 天前

There is. And no it can't predict very far ahead. You see these tendencies to go "left" and "down-right" zigzag? These are trained words. check jevs chess demo - you'd see without intelligence you can't do much with a simple classifier

Siddharth Jain 的头像
Siddharth Jain7 天前

Jax is a machine learning library? Is it the same or you named it the same?

Brjan | AI Builder 的头像
Brjan | AI Builder7 天前

what specific benchmarks did you use to compare Jax against Jev and Laya?

Shez Malik 的头像
Shez Malik7 天前

proud of u son

相关视频