Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Benchmark idea: Who can use fewest lines of code to do the same thing? In this test, Opus 5.5 uses ~half the lines of code that of Astra and the physics/visual quality isn't noticeably worse. I am trying to approximate how 'elegant' the code is - something that many...

32,183 görüntüleme • 2 gün önce •via X (Twitter)

20 Yorum

Ryan Sael profil fotoğrafı
Ryan Sael2 gün önce

I was thinking the same! It's crazy that these files have less than 200kb each. I'm open sourcing these:

Peter Gostev (in SF @ DevDay) profil fotoğrafı
Peter Gostev (in SF @ DevDay)2 gün önce

This is the breakdown of the lines of code - Astra routinely spent lots of LoCs on the details of collision, so it isn't that it just put lots of comments in. Which code is 'better' hard for me to judge, but if there's interest I can open source too.

Peter Gostev (in SF @ DevDay) profil fotoğrafı
Peter Gostev (in SF @ DevDay)2 gün önce

Another important concept that tokens != lines of code. It should be OK to spend lots of tokens and end up with less code. If you look at 'Max' reasoning in particular, Opus spent 175k tokens, while Astra spent 25k tokens, but it ended up with close to half the LoC. While obviously more expensive, I'm interested in this mechanic of 'thinking hard to solve problem more elegantly' - because this is what humans tend to do more.

Peter Gostev (in SF @ DevDay) profil fotoğrafı
Peter Gostev (in SF @ DevDay)2 gün önce

The way this benchmark works is that we have a reference design (i.e. screenshot), specification, request for realistic physics etc, and importantly - fewest lines of code to implement. I have tried a few variations of this, but this shape feels right.

This is Greg profil fotoğrafı
This is Greg2 gün önce

what would make fewest lines better though? code is so incredibly cheap.

Peter Gostev (in SF @ DevDay) profil fotoğrafı
Peter Gostev (in SF @ DevDay)2 gün önce

Because you generally don't want to end up with enormous code bases that are impossible to maintain. More code could also mean that it 'hacks' each individual case instead of eg designing a proper data schema. I had apps that ended up with 30k SQL code when 3k was sufficient

This is Greg profil fotoğrafı
This is Greg2 gün önce

fair enough. i actually use to put constraints on the models for that reason 9-12 months ago because i saw code bases grow, but now the models are so good that it almost doesn't matter as much. i threw opus 5.5 at an old codebase and told it to review everything and find anything it wanted to fix/improve. it replaced ~20k lines of code and found so many issues but also added back a fair amount

Pulso IA | Noticias en español profil fotoğrafı
Pulso IA | Noticias en español2 gün önce

I'd add a second round: ask each model to change one physics rule, then rerun the same collision tests. Fewer lines would be much more convincing if the smaller version is also easier to change without breaking existing behavior.

Adrian Rangel profil fotoğrafı
Adrian Rangel2 gün önce

please also include comments in the metric. I hate when models are super verbose. 10 loc + 100 lines of comments

Peter Gostev (in SF @ DevDay) profil fotoğrafı
Peter Gostev (in SF @ DevDay)2 gün önce

There weren't that many in this one, I like about 12-14 lines each

Vedant Padwal profil fotoğrafı
Vedant Padwal2 gün önce

This idea already exists and is called code golf, I created a benchmark for this.

Hope profil fotoğrafı
Hope2 gün önce

it actually is worse as the blocks are clipping/overlapping and the physics model is not as good - compare these clips in slow motion

Gil MD profil fotoğrafı
Gil MD2 gün önce

Astra looks better quality than opus here on finer look at the details

Conan Reis 🇨🇦 profil fotoğrafı
Conan Reis 🇨🇦2 gün önce

Interesting in principle - though lines of code is arbitrary. As a programming language designer number of expressions would be better - independent of LOC. Still probably needs even more sophistication for a good metric. Labs might do this just for token efficiency.

NaN goto profil fotoğrafı
NaN goto2 gün önce

You should use assembly instructions as the comparison not LoC.

ρ:ɡeon profil fotoğrafı
ρ:ɡeon2 gün önce

code golf bench

Gregor profil fotoğrafı
Gregor2 gün önce

When I collapsed a Flutter data layer from 200 to 80 lines, bugs took twice as long to trace. The compressed code looked elegant. The stack traces disagreed.

woody lee profil fotoğrafı
woody lee2 gün önce

@scaling01 Do you know enough about the physics to judge design trade offs in the physics engine though? Without that I think it would be hard to judge.

Francisco profil fotoğrafı
Francisco2 gün önce

Never heard of code golf?

Juaki profil fotoğrafı
Juaki2 gün önce

Astra is more aggressive when it comes to debugging, and that ultimately carries over into the code it writes. Sometimes it can also be pretty stubborn about adding redundancies and fallback/escape paths to force an outcome instead of simply stopping the flow when it should. Claude tends to just make things work. It generally handles control flow and lifecycles better, with fewer unnecessary workarounds. Both produce excellent code, but their approaches are noticeably different.

Benzer Videolar

AI is changing the software engineering craft. Anders Hejlsberg (Anders Hejlsberg) - creator of C#, TypeScript and industry legend - on why code review needs to get more enjoyable in response: #1 - AI is shifting the craft from writing code, to reviewing code: "In a sense, we're all turning into project managers. We can have an army of junior programmers, called agents, that will just spit out reams of code but someone's got to have the big picture and review all of that. And so, increasingly, our craft is going from one of writing the code, to one of reviewing the code and building the architecture of the code and overseeing the work. It's a different kind of craft. It's a different kind of enjoyment. I've always liked writing the code. To me that was the fulfilling part, seeing it work. In a way, AI robs a little bit of that, because I am less interested in reviewing code." #2 - The code review experience should be improved: "I think we could also make the process of reviewing code much more interesting than it is today. I mean, today, you see a list of diffs in alphabetical order and now it's up to you to make heads or tails of it. There are more pedagogical ways of presenting that. And you could have commentary generated by the AI that tells you what the changes are and whatever, and then tries to guide you along. So that symbiotic relationship, I think we need to work on that more and to keep the enjoyment in there."

The Pragmatic Engineer

39,073 görüntüleme • 4 ay önce