正在加载视频...

视频加载失败

We now run evals to help you know what AI models perform best to write Nuxt code.

28,195 次观看 • 7 个月前 •via X (Twitter)

13 条评论

Matt Maribojoc 的头像
Matt Maribojoc7 个月前

lets go. curious if nuxt mcp or the skills actually help with this.

Darek Gusto 的头像
Darek Gusto7 个月前

That's such a good idea. But no open weight models = no evals.

Mathieu 的头像
Mathieu7 个月前

Hooo, would love to have this for nuxt ui !

Estéban 的头像
Estéban7 个月前

Great idea! (Side note, it's impossible to access this page using the search and there's not link to it from the mobile navigation)

Shiyam Kashfiq 的头像
Shiyam Kashfiq7 个月前

Smart move. Clear signal on how to pick AI helpers

Serhii Chernenko 的头像
Serhii Chernenko7 个月前

I can have doubts about it. An hour ago I wrote a custom GET fetching logic directly in a component. Opus 4.6 didn't use `await` before `useAsyncData`. It didn't know about the `watch` option. Implemented everything with $fetch and overhead. It's was a simple task that could be handled by this doc: And yes. I have installed with the skills directory. I even have some custom notes to improve DX for it but it still fails. Meanwhile, I didn't have any stupid issues with Codex 5.3

Lukas Trumm 的头像
Lukas Trumm7 个月前

Awesome! My experience is that for opinionated rules, its still better to mostly include the rules and fine tune them for better results. Might be short, models understand. Here is my own take on evals (focusing on opinionated Vue rules and Claude Code).

andrej.fidel 的头像
andrej.fidel7 个月前

Woud love to see a gpt 5.2 banchmark, I've been having better luck with vanilla 5.2 than 5.3 codex on frameworks like Laravel, probably because the framework abstracts a lot of concepts into something resembling human speech rather than algorithms or competitive coding style. Haven't done deep dives with Nuxt yet though.

Stefan Galescu 的头像
Stefan Galescu7 个月前

I’d be curious to see if 5.3 Codex performs better using @opencode

Max Gfeller 的头像
Max Gfeller7 个月前

Love this idea! Was thinking of doing something like that for a thing we’re currently working on.

👋Hey Alex 的头像
👋Hey Alex7 个月前

@hugorcd The overall percentage on a small number of evals is a bit confusing. It seems there is a huge difference between Opus and Codex. But when you look into it it is a matter of a couple of tests. I would prefer to see 24/25 and 22/25.

JD Solanki 的头像
JD Solanki7 个月前

Thanks

Ismael 的头像
Ismael7 个月前

Fantastic, I've been using codex 5.3 lately, but looks I shoul've been using opus 4.6!

相关视频