Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

We now run evals to help you know what AI models perform best to write Nuxt code.

28,195 Aufrufe • vor 7 Monaten •via X (Twitter)

13 Kommentare

Profilbild von Matt Maribojoc
Matt Maribojocvor 7 Monaten

lets go. curious if nuxt mcp or the skills actually help with this.

Profilbild von Darek Gusto
Darek Gustovor 7 Monaten

That's such a good idea. But no open weight models = no evals.

Profilbild von Mathieu
Mathieuvor 7 Monaten

Hooo, would love to have this for nuxt ui !

Profilbild von Estéban
Estébanvor 7 Monaten

Great idea! (Side note, it's impossible to access this page using the search and there's not link to it from the mobile navigation)

Profilbild von Shiyam Kashfiq
Shiyam Kashfiqvor 7 Monaten

Smart move. Clear signal on how to pick AI helpers

Profilbild von Serhii Chernenko
Serhii Chernenkovor 7 Monaten

I can have doubts about it. An hour ago I wrote a custom GET fetching logic directly in a component. Opus 4.6 didn't use `await` before `useAsyncData`. It didn't know about the `watch` option. Implemented everything with $fetch and overhead. It's was a simple task that could be handled by this doc: And yes. I have installed with the skills directory. I even have some custom notes to improve DX for it but it still fails. Meanwhile, I didn't have any stupid issues with Codex 5.3

Profilbild von Lukas Trumm
Lukas Trummvor 7 Monaten

Awesome! My experience is that for opinionated rules, its still better to mostly include the rules and fine tune them for better results. Might be short, models understand. Here is my own take on evals (focusing on opinionated Vue rules and Claude Code).

Profilbild von andrej.fidel
andrej.fidelvor 7 Monaten

Woud love to see a gpt 5.2 banchmark, I've been having better luck with vanilla 5.2 than 5.3 codex on frameworks like Laravel, probably because the framework abstracts a lot of concepts into something resembling human speech rather than algorithms or competitive coding style. Haven't done deep dives with Nuxt yet though.

Profilbild von Stefan Galescu
Stefan Galescuvor 7 Monaten

I’d be curious to see if 5.3 Codex performs better using @opencode

Profilbild von Max Gfeller
Max Gfellervor 7 Monaten

Love this idea! Was thinking of doing something like that for a thing we’re currently working on.

Profilbild von 👋Hey Alex
👋Hey Alexvor 7 Monaten

@hugorcd The overall percentage on a small number of evals is a bit confusing. It seems there is a huge difference between Opus and Codex. But when you look into it it is a matter of a couple of tests. I would prefer to see 24/25 and 22/25.

Profilbild von JD Solanki
JD Solankivor 7 Monaten

Thanks

Profilbild von Ismael
Ismaelvor 7 Monaten

Fantastic, I've been using codex 5.3 lately, but looks I shoul've been using opus 4.6!

Ähnliche Videos