Video wird geladen...
Video konnte nicht geladen werden
Benchmark idea: Who can use fewest lines of code to do the same thing? In this test, Opus 5.5 uses ~half the lines of code that of Astra and the physics/visual quality isn't noticeably worse. I am trying to approximate how 'elegant' the code is - something that many... show more
32,183 Aufrufe • vor 3 Tagen •via X (Twitter)
20 Kommentare

I was thinking the same! It's crazy that these files have less than 200kb each. I'm open sourcing these:

This is the breakdown of the lines of code - Astra routinely spent lots of LoCs on the details of collision, so it isn't that it just put lots of comments in. Which code is 'better' hard for me to judge, but if there's interest I can open source too.

Another important concept that tokens != lines of code. It should be OK to spend lots of tokens and end up with less code. If you look at 'Max' reasoning in particular, Opus spent 175k tokens, while Astra spent 25k tokens, but it ended up with close to half the LoC. While obviously more expensive, I'm interested in this mechanic of 'thinking hard to solve problem more elegantly' - because this is what humans tend to do more.

The way this benchmark works is that we have a reference design (i.e. screenshot), specification, request for realistic physics etc, and importantly - fewest lines of code to implement. I have tried a few variations of this, but this shape feels right.

what would make fewest lines better though? code is so incredibly cheap.

Because you generally don't want to end up with enormous code bases that are impossible to maintain. More code could also mean that it 'hacks' each individual case instead of eg designing a proper data schema. I had apps that ended up with 30k SQL code when 3k was sufficient

fair enough. i actually use to put constraints on the models for that reason 9-12 months ago because i saw code bases grow, but now the models are so good that it almost doesn't matter as much. i threw opus 5.5 at an old codebase and told it to review everything and find anything it wanted to fix/improve. it replaced ~20k lines of code and found so many issues but also added back a fair amount

I'd add a second round: ask each model to change one physics rule, then rerun the same collision tests. Fewer lines would be much more convincing if the smaller version is also easier to change without breaking existing behavior.

please also include comments in the metric. I hate when models are super verbose. 10 loc + 100 lines of comments

There weren't that many in this one, I like about 12-14 lines each

This idea already exists and is called code golf, I created a benchmark for this.

it actually is worse as the blocks are clipping/overlapping and the physics model is not as good - compare these clips in slow motion

Astra looks better quality than opus here on finer look at the details

Interesting in principle - though lines of code is arbitrary. As a programming language designer number of expressions would be better - independent of LOC. Still probably needs even more sophistication for a good metric. Labs might do this just for token efficiency.

You should use assembly instructions as the comparison not LoC.

code golf bench

When I collapsed a Flutter data layer from 200 to 80 lines, bugs took twice as long to trace. The compressed code looked elegant. The stack traces disagreed.

@scaling01 Do you know enough about the physics to judge design trade offs in the physics engine though? Without that I think it would be hard to judge.

Never heard of code golf?

Astra is more aggressive when it comes to debugging, and that ultimately carries over into the code it writes. Sometimes it can also be pretty stubborn about adding redundancies and fallback/escape paths to force an outcome instead of simply stopping the flow when it should. Claude tends to just make things work. It generally handles control flow and lifecycles better, with fewer unnecessary workarounds. Both produce excellent code, but their approaches are noticeably different.

