Загрузка видео...

Не удалось загрузить видео

На главную

We now run evals to help you know what AI models perform best to write Nuxt code.

28,195 просмотров • 7 месяцев назад •via X (Twitter)

Комментарии: 13

Фото профиля Matt Maribojoc
Matt Maribojoc7 месяцев назад

lets go. curious if nuxt mcp or the skills actually help with this.

Фото профиля Darek Gusto
Darek Gusto7 месяцев назад

That's such a good idea. But no open weight models = no evals.

Фото профиля Mathieu
Mathieu7 месяцев назад

Hooo, would love to have this for nuxt ui !

Фото профиля Estéban
Estéban7 месяцев назад

Great idea! (Side note, it's impossible to access this page using the search and there's not link to it from the mobile navigation)

Фото профиля Shiyam Kashfiq
Shiyam Kashfiq7 месяцев назад

Smart move. Clear signal on how to pick AI helpers

Фото профиля Serhii Chernenko
Serhii Chernenko7 месяцев назад

I can have doubts about it. An hour ago I wrote a custom GET fetching logic directly in a component. Opus 4.6 didn't use `await` before `useAsyncData`. It didn't know about the `watch` option. Implemented everything with $fetch and overhead. It's was a simple task that could be handled by this doc: And yes. I have installed with the skills directory. I even have some custom notes to improve DX for it but it still fails. Meanwhile, I didn't have any stupid issues with Codex 5.3

Фото профиля Lukas Trumm
Lukas Trumm7 месяцев назад

Awesome! My experience is that for opinionated rules, its still better to mostly include the rules and fine tune them for better results. Might be short, models understand. Here is my own take on evals (focusing on opinionated Vue rules and Claude Code).

Фото профиля andrej.fidel
andrej.fidel7 месяцев назад

Woud love to see a gpt 5.2 banchmark, I've been having better luck with vanilla 5.2 than 5.3 codex on frameworks like Laravel, probably because the framework abstracts a lot of concepts into something resembling human speech rather than algorithms or competitive coding style. Haven't done deep dives with Nuxt yet though.

Фото профиля Stefan Galescu
Stefan Galescu7 месяцев назад

I’d be curious to see if 5.3 Codex performs better using @opencode

Фото профиля Max Gfeller
Max Gfeller7 месяцев назад

Love this idea! Was thinking of doing something like that for a thing we’re currently working on.

Фото профиля 👋Hey Alex
👋Hey Alex7 месяцев назад

@hugorcd The overall percentage on a small number of evals is a bit confusing. It seems there is a huge difference between Opus and Codex. But when you look into it it is a matter of a couple of tests. I would prefer to see 24/25 and 22/25.

Фото профиля JD Solanki
JD Solanki7 месяцев назад

Thanks

Фото профиля Ismael
Ismael7 месяцев назад

Fantastic, I've been using codex 5.3 lately, but looks I shoul've been using opus 4.6!

Похожие видео