Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

We now run evals to help you know what AI models perform best to write Nuxt code.

28,195 görüntüleme • 7 ay önce •via X (Twitter)

13 Yorum

Matt Maribojoc profil fotoğrafı
Matt Maribojoc7 ay önce

lets go. curious if nuxt mcp or the skills actually help with this.

Darek Gusto profil fotoğrafı
Darek Gusto7 ay önce

That's such a good idea. But no open weight models = no evals.

Mathieu profil fotoğrafı
Mathieu7 ay önce

Hooo, would love to have this for nuxt ui !

Estéban profil fotoğrafı
Estéban7 ay önce

Great idea! (Side note, it's impossible to access this page using the search and there's not link to it from the mobile navigation)

Shiyam Kashfiq profil fotoğrafı
Shiyam Kashfiq7 ay önce

Smart move. Clear signal on how to pick AI helpers

Serhii Chernenko profil fotoğrafı
Serhii Chernenko7 ay önce

I can have doubts about it. An hour ago I wrote a custom GET fetching logic directly in a component. Opus 4.6 didn't use `await` before `useAsyncData`. It didn't know about the `watch` option. Implemented everything with $fetch and overhead. It's was a simple task that could be handled by this doc: And yes. I have installed with the skills directory. I even have some custom notes to improve DX for it but it still fails. Meanwhile, I didn't have any stupid issues with Codex 5.3

Lukas Trumm profil fotoğrafı
Lukas Trumm7 ay önce

Awesome! My experience is that for opinionated rules, its still better to mostly include the rules and fine tune them for better results. Might be short, models understand. Here is my own take on evals (focusing on opinionated Vue rules and Claude Code).

andrej.fidel profil fotoğrafı
andrej.fidel7 ay önce

Woud love to see a gpt 5.2 banchmark, I've been having better luck with vanilla 5.2 than 5.3 codex on frameworks like Laravel, probably because the framework abstracts a lot of concepts into something resembling human speech rather than algorithms or competitive coding style. Haven't done deep dives with Nuxt yet though.

Stefan Galescu profil fotoğrafı
Stefan Galescu7 ay önce

I’d be curious to see if 5.3 Codex performs better using @opencode

Max Gfeller profil fotoğrafı
Max Gfeller7 ay önce

Love this idea! Was thinking of doing something like that for a thing we’re currently working on.

👋Hey Alex profil fotoğrafı
👋Hey Alex7 ay önce

@hugorcd The overall percentage on a small number of evals is a bit confusing. It seems there is a huge difference between Opus and Codex. But when you look into it it is a matter of a couple of tests. I would prefer to see 24/25 and 22/25.

JD Solanki profil fotoğrafı
JD Solanki7 ay önce

Thanks

Ismael profil fotoğrafı
Ismael7 ay önce

Fantastic, I've been using codex 5.3 lately, but looks I shoul've been using opus 4.6!

Benzer Videolar