Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Benchmark idea: Who can use fewest lines of code to do the same thing? In this test, Opus 5.5 uses ~half the lines of code that of Astra and the physics/visual quality isn't noticeably worse. I am trying to approximate how 'elegant' the code is - something that many...

32,183 Aufrufe • vor 3 Tagen •via X (Twitter)

20 Kommentare

Profilbild von Ryan Sael
Ryan Saelvor 3 Tagen

I was thinking the same! It's crazy that these files have less than 200kb each. I'm open sourcing these:

Profilbild von Peter Gostev (in SF @ DevDay)
Peter Gostev (in SF @ DevDay)vor 3 Tagen

This is the breakdown of the lines of code - Astra routinely spent lots of LoCs on the details of collision, so it isn't that it just put lots of comments in. Which code is 'better' hard for me to judge, but if there's interest I can open source too.

Profilbild von Peter Gostev (in SF @ DevDay)
Peter Gostev (in SF @ DevDay)vor 3 Tagen

Another important concept that tokens != lines of code. It should be OK to spend lots of tokens and end up with less code. If you look at 'Max' reasoning in particular, Opus spent 175k tokens, while Astra spent 25k tokens, but it ended up with close to half the LoC. While obviously more expensive, I'm interested in this mechanic of 'thinking hard to solve problem more elegantly' - because this is what humans tend to do more.

Profilbild von Peter Gostev (in SF @ DevDay)
Peter Gostev (in SF @ DevDay)vor 3 Tagen

The way this benchmark works is that we have a reference design (i.e. screenshot), specification, request for realistic physics etc, and importantly - fewest lines of code to implement. I have tried a few variations of this, but this shape feels right.

Profilbild von This is Greg
This is Gregvor 3 Tagen

what would make fewest lines better though? code is so incredibly cheap.

Profilbild von Peter Gostev (in SF @ DevDay)
Peter Gostev (in SF @ DevDay)vor 3 Tagen

Because you generally don't want to end up with enormous code bases that are impossible to maintain. More code could also mean that it 'hacks' each individual case instead of eg designing a proper data schema. I had apps that ended up with 30k SQL code when 3k was sufficient

Profilbild von This is Greg
This is Gregvor 3 Tagen

fair enough. i actually use to put constraints on the models for that reason 9-12 months ago because i saw code bases grow, but now the models are so good that it almost doesn't matter as much. i threw opus 5.5 at an old codebase and told it to review everything and find anything it wanted to fix/improve. it replaced ~20k lines of code and found so many issues but also added back a fair amount

Profilbild von Pulso IA | Noticias en español
Pulso IA | Noticias en españolvor 3 Tagen

I'd add a second round: ask each model to change one physics rule, then rerun the same collision tests. Fewer lines would be much more convincing if the smaller version is also easier to change without breaking existing behavior.

Profilbild von Adrian Rangel
Adrian Rangelvor 2 Tagen

please also include comments in the metric. I hate when models are super verbose. 10 loc + 100 lines of comments

Profilbild von Peter Gostev (in SF @ DevDay)
Peter Gostev (in SF @ DevDay)vor 2 Tagen

There weren't that many in this one, I like about 12-14 lines each

Profilbild von Vedant Padwal
Vedant Padwalvor 2 Tagen

This idea already exists and is called code golf, I created a benchmark for this.

Profilbild von Hope
Hopevor 2 Tagen

it actually is worse as the blocks are clipping/overlapping and the physics model is not as good - compare these clips in slow motion

Profilbild von Gil MD
Gil MDvor 2 Tagen

Astra looks better quality than opus here on finer look at the details

Profilbild von Conan Reis 🇨🇦
Conan Reis 🇨🇦vor 2 Tagen

Interesting in principle - though lines of code is arbitrary. As a programming language designer number of expressions would be better - independent of LOC. Still probably needs even more sophistication for a good metric. Labs might do this just for token efficiency.

Profilbild von NaN goto
NaN gotovor 2 Tagen

You should use assembly instructions as the comparison not LoC.

Profilbild von ρ:ɡeon
ρ:ɡeonvor 2 Tagen

code golf bench

Profilbild von Gregor
Gregorvor 2 Tagen

When I collapsed a Flutter data layer from 200 to 80 lines, bugs took twice as long to trace. The compressed code looked elegant. The stack traces disagreed.

Profilbild von woody lee
woody leevor 2 Tagen

@scaling01 Do you know enough about the physics to judge design trade offs in the physics engine though? Without that I think it would be hard to judge.

Profilbild von Francisco
Franciscovor 2 Tagen

Never heard of code golf?

Profilbild von Juaki
Juakivor 2 Tagen

Astra is more aggressive when it comes to debugging, and that ultimately carries over into the code it writes. Sometimes it can also be pretty stubborn about adding redundancies and fallback/escape paths to force an outcome instead of simply stopping the flow when it should. Claude tends to just make things work. It generally handles control flow and lifecycles better, with fewer unnecessary workarounds. Both produce excellent code, but their approaches are noticeably different.

Ähnliche Videos

AI is changing the software engineering craft. Anders Hejlsberg (Anders Hejlsberg) - creator of C#, TypeScript and industry legend - on why code review needs to get more enjoyable in response: #1 - AI is shifting the craft from writing code, to reviewing code: "In a sense, we're all turning into project managers. We can have an army of junior programmers, called agents, that will just spit out reams of code but someone's got to have the big picture and review all of that. And so, increasingly, our craft is going from one of writing the code, to one of reviewing the code and building the architecture of the code and overseeing the work. It's a different kind of craft. It's a different kind of enjoyment. I've always liked writing the code. To me that was the fulfilling part, seeing it work. In a way, AI robs a little bit of that, because I am less interested in reviewing code." #2 - The code review experience should be improved: "I think we could also make the process of reviewing code much more interesting than it is today. I mean, today, you see a list of diffs in alphabetical order and now it's up to you to make heads or tails of it. There are more pedagogical ways of presenting that. And you could have commentary generated by the AI that tells you what the changes are and whatever, and then tries to guide you along. So that symbiotic relationship, I think we need to work on that more and to keep the enjoyment in there."

The Pragmatic Engineer

39,073 Aufrufe • vor 4 Monaten