Loading video...

Video Failed to Load

Go Home

We now run evals to help you know what AI models perform best to write Nuxt code.

28,195 views • 7 months ago •via X (Twitter)

13 Comments

Matt Maribojoc's profile picture
Matt Maribojoc7 months ago

lets go. curious if nuxt mcp or the skills actually help with this.

Darek Gusto's profile picture
Darek Gusto7 months ago

That's such a good idea. But no open weight models = no evals.

Mathieu's profile picture
Mathieu7 months ago

Hooo, would love to have this for nuxt ui !

Estéban's profile picture
Estéban7 months ago

Great idea! (Side note, it's impossible to access this page using the search and there's not link to it from the mobile navigation)

Shiyam Kashfiq's profile picture
Shiyam Kashfiq7 months ago

Smart move. Clear signal on how to pick AI helpers

Serhii Chernenko's profile picture
Serhii Chernenko7 months ago

I can have doubts about it. An hour ago I wrote a custom GET fetching logic directly in a component. Opus 4.6 didn't use `await` before `useAsyncData`. It didn't know about the `watch` option. Implemented everything with $fetch and overhead. It's was a simple task that could be handled by this doc: And yes. I have installed with the skills directory. I even have some custom notes to improve DX for it but it still fails. Meanwhile, I didn't have any stupid issues with Codex 5.3

Lukas Trumm's profile picture
Lukas Trumm7 months ago

Awesome! My experience is that for opinionated rules, its still better to mostly include the rules and fine tune them for better results. Might be short, models understand. Here is my own take on evals (focusing on opinionated Vue rules and Claude Code).

andrej.fidel's profile picture
andrej.fidel7 months ago

Woud love to see a gpt 5.2 banchmark, I've been having better luck with vanilla 5.2 than 5.3 codex on frameworks like Laravel, probably because the framework abstracts a lot of concepts into something resembling human speech rather than algorithms or competitive coding style. Haven't done deep dives with Nuxt yet though.

Stefan Galescu's profile picture
Stefan Galescu7 months ago

I’d be curious to see if 5.3 Codex performs better using @opencode

Max Gfeller's profile picture
Max Gfeller7 months ago

Love this idea! Was thinking of doing something like that for a thing we’re currently working on.

👋Hey Alex's profile picture
👋Hey Alex7 months ago

@hugorcd The overall percentage on a small number of evals is a bit confusing. It seems there is a huge difference between Opus and Codex. But when you look into it it is a matter of a couple of tests. I would prefer to see 24/25 and 22/25.

JD Solanki's profile picture
JD Solanki7 months ago

Thanks

Ismael's profile picture
Ismael7 months ago

Fantastic, I've been using codex 5.3 lately, but looks I shoul've been using opus 4.6!

Related Videos