pilvar (Philippe Dourassov)'s banner
pilvar (Philippe Dourassov)'s profile picture

pilvar (Philippe Dourassov)

@pilvar2222,948 subscribers

AI Pentest Lead @AikidoSecurity

Shorts

Holy moly: GLM-5.3 got much better in cybersecurity since our pre-release evaluation with Z.ai. It now matches GPT-5.6-Sol on our cybersecurity benchmark at 0.4x the cost 🤯 - At pass@1: it went from 60.4% to 65.6% CVEs rediscovered, crushing every other open model on one-shot tasks - At pass@3: it did 75% -> 78.1%, matching GPT-5.6-Sol - Its precision remained stable, reporting fewer false positives than DeepSeek models The performance increase comes from a behavioral change: the new version is more persistent. It tends to run longer, and had a ~43% reasoning tokens increase. But the performance upgrade is worth that additional cost. 1/3 🧵

Holy moly: GLM-5.3 got much better in cybersecurity since our pre-release evaluation with Z.ai. It now matches GPT-5.6-Sol on our cybersecurity benchmark at 0.4x the cost 🤯 - At pass@1: it went from 60.4% to 65.6% CVEs rediscovered, crushing every other open model on one-shot tasks - At pass@3: it did 75% -> 78.1%, matching GPT-5.6-Sol - Its precision remained stable, reporting fewer false positives than DeepSeek models The performance increase comes from a behavioral change: the new version is more persistent. It tends to run longer, and had a ~43% reasoning tokens increase. But the performance upgrade is worth that additional cost. 1/3 🧵

22,246 görüntüleme

Videos

Daha fazla içerik yok.