Video yükleniyor...
Video Yüklenemedi
Still didn't attempt any quality testing, but let's say that GPT 5.6 Sol was able almost unassisted to write this Laguna S2.1 inference implementation following the other two models in DwarfStar, and it is very fast, 50 t/s generation, very fast prefill as well.
43,676 görüntüleme • 2 ay önce •via X (Twitter)
22 Yorum

50 t/s is nice, but Sol writing Laguna support from two neighboring implementations is the wilder result. Inference backends are becoming executable architecture docs.

I hope your version will change my mind, but for my use case, Laguna has been a disaster so far!

I'm also seeing odd things for the Q4 quants.

Same experience with Q4.

That’s as fast as gpt-5.6-luna api !

50 t/s is a marketing number until the quality pass happens. fast inference on wrong output is just a quicker way to be wrong.

Sol is extremely talented and helpful goblin. Substantially less mistakes and bs than Fable.

When are you going to implement --n-cpu-moe ? Can GPT 5.6-Sol implement it ?

Impressive work! I understand this is Q4: how does it compare to the other local models?

Have you had a chance to benchmark the models on DwarfStar for both speed and reasoning?

50 t/s is nice, but the real eval is the cleanup diff. How much of Sol’s implementation survived your review unchanged?

Do you think it's still useful to learn how to write complex kernels?

Can you run the tests ? On mine it's worse then flash V4.

what hardware ?

m5 max 128gb

united open agents

Heard mixed things about the quality. Coding is great. Other stuff depends.

how is the decline with context depth increase - in absolute, and compared to ds4f?

correctness in inference code was never the hard part. models have written plausible attention loops for a year now. 50 t/s with fast prefill on an m5 max is a different animal. kv cache layout, the 4-bit paths, kernels that don't stall on metal. the stuff you only get right after sitting with a profiler for a week. if sol pulled that off almost unprompted, then the last skill i figured still needed a human is going too. hardware-aware perf work. the unglamorous profiling nobody wants to touch.

sol cooked up its own rocket engine

Give a shot to benchlocal-cli

Would love to hear about how you go about prompting these days.
