Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Still didn't attempt any quality testing, but let's say that GPT 5.6 Sol was able almost unassisted to write this Laguna S2.1 inference implementation following the other two models in DwarfStar, and it is very fast, 50 t/s generation, very fast prefill as well.

43,676 görüntüleme • 2 ay önce •via X (Twitter)

22 Yorum

Sakura Yuki profil fotoğrafı
Sakura Yuki2 ay önce

50 t/s is nice, but Sol writing Laguna support from two neighboring implementations is the wilder result. Inference backends are becoming executable architecture docs.

Giovanni profil fotoğrafı
Giovanni2 ay önce

I hope your version will change my mind, but for my use case, Laguna has been a disaster so far!

antirez profil fotoğrafı
antirez2 ay önce

I'm also seeing odd things for the Q4 quants.

Giovanni profil fotoğrafı
Giovanni2 ay önce

Same experience with Q4.

Shubham Arora profil fotoğrafı
Shubham Arora2 ay önce

That’s as fast as gpt-5.6-luna api !

Michał Piszczek profil fotoğrafı
Michał Piszczek2 ay önce

50 t/s is a marketing number until the quality pass happens. fast inference on wrong output is just a quicker way to be wrong.

BitXChange profil fotoğrafı
BitXChange2 ay önce

Sol is extremely talented and helpful goblin. Substantially less mistakes and bs than Fable.

inside inside profil fotoğrafı
inside inside2 ay önce

When are you going to implement --n-cpu-moe ? Can GPT 5.6-Sol implement it ?

Ɠuglièlmo "𝚐𝚠𝚣" Ɲígri profil fotoğrafı
Ɠuglièlmo "𝚐𝚠𝚣" Ɲígri2 ay önce

Impressive work! I understand this is Q4: how does it compare to the other local models?

Francesco Marasco 🇮🇹🇪🇺 profil fotoğrafı
Francesco Marasco 🇮🇹🇪🇺2 ay önce

Have you had a chance to benchmark the models on DwarfStar for both speed and reasoning?

LeetLLM.com profil fotoğrafı
LeetLLM.com2 ay önce

50 t/s is nice, but the real eval is the cleanup diff. How much of Sol’s implementation survived your review unchanged?

clainstone profil fotoğrafı
clainstone2 ay önce

Do you think it's still useful to learn how to write complex kernels?

pls seed profil fotoğrafı
pls seed2 ay önce

Can you run the tests ? On mine it's worse then flash V4.

Diego Carlino profil fotoğrafı
Diego Carlino2 ay önce

what hardware ?

antirez profil fotoğrafı
antirez2 ay önce

m5 max 128gb

Angel profil fotoğrafı
Angel2 ay önce

united open agents

Dustin Ogle profil fotoğrafı
Dustin Ogle2 ay önce

Heard mixed things about the quality. Coding is great. Other stuff depends.

Ljubomir Josifovski profil fotoğrafı
Ljubomir Josifovski2 ay önce

how is the decline with context depth increase - in absolute, and compared to ds4f?

Daksh Trehan profil fotoğrafı
Daksh Trehan2 ay önce

correctness in inference code was never the hard part. models have written plausible attention loops for a year now. 50 t/s with fast prefill on an m5 max is a different animal. kv cache layout, the 4-bit paths, kernels that don't stall on metal. the stuff you only get right after sitting with a profiler for a week. if sol pulled that off almost unprompted, then the last skill i figured still needed a human is going too. hardware-aware perf work. the unglamorous profiling nobody wants to touch.

shaped profil fotoğrafı
shaped2 ay önce

sol cooked up its own rocket engine

noname profil fotoğrafı
noname2 ay önce

Give a shot to benchlocal-cli

Jedd Haberstro profil fotoğrafı
Jedd Haberstro2 ay önce

Would love to hear about how you go about prompting these days.

Benzer Videolar

GPT-5.6 vs GPT-5.5 on my custom spaceship prompt. I gave both models the exact same custom prompt. This is also the same prompt I previously gave to Fable 5. For context, GPT-5.6 Pro worked for 87 minutes, while GPT-5.5 Extra High worked for 34 minutes and 42 seconds. As I’ve said before, based on great authority GPT-5.6 will be an incremental/soldi improvement over GPT-5.5, not a “Fable killer.” My rough expectation has been that it would trade blows with Fable 5 on some benchmarks, maybe win around half depending on the category, but not clearly surpass it overall. And again fable five will have bigger model smell, but this was expected. After testing this coding output, that view feels pretty accurate. GPT-5.6 is clearly better than GPT-5.5 in several visual areas. The lighting, shading, chairs, object details, and exterior of the spaceship looked noticeably stronger. The scene was also easier to test. I do want to give GPT-5.5 credit though. It built out the rooms much much better and the planets looked better than GPT-5.6’s. It was also interesting that both GPT-5.5 and GPT-5.6 produced better-looking planets than Fable 5 in this specific test. The downside with GPT-5.5 was stability. The game was much glitchier and harder to test compared to GPT-5.6. But when it comes to the core of the demo, which is the spaceship itself, Fable 5 still beat both models pretty comfortably. GPT-5.6 is impressive, but from this test, it looks exactly like what I expected which was a meaningful incremental improvement over GPT-5.5, at least for indie game demos, but not something that replaces Fable 5. In collaboration with Chetaslua

Chris

250,919 görüntüleme • 3 ay önce