pilvar (Philippe Dourassov)'s banner
pilvar (Philippe Dourassov)'s profile picture

pilvar (Philippe Dourassov)

@pilvar222 • 2,948 subscribers

AI Pentest Lead @AikidoSecurity

Shorts

HOLY MOLY: Aikido got GPT-6 Astra in advance to run it on our Cybersecurity benchmark, it crushed EVERY other model! - At pass@3 it rediscovered 29/32 CVEs, the highest recall we've ever recorded and 4 more than GPT-5.6-Sol - Even at pass@1, it has 75% recall. The model is VERY consistent - The performance however come at a high price (literally). The three runs cost us almost $4,000 Astra is now the #1 model on the benchmark Debarshi and I built, and by a LOT 1/3 🧵

HOLY MOLY: Aikido got GPT-6 Astra in advance to run it on our Cybersecurity benchmark, it crushed EVERY other model! - At pass@3 it rediscovered 29/32 CVEs, the highest recall we've ever recorded and 4 more than GPT-5.6-Sol - Even at pass@1, it has 75% recall. The model is VERY consistent - The performance however come at a high price (literally). The three runs cost us almost $4,000 Astra is now the #1 model on the benchmark Debarshi and I built, and by a LOT 1/3 🧵

49,224 просмотров

Holy moly: GLM-5.3 got much better in cybersecurity since our pre-release evaluation with Z.ai. It now matches GPT-5.6-Sol on our cybersecurity benchmark at 0.4x the cost 🤯 - At pass@1: it went from 60.4% to 65.6% CVEs rediscovered, crushing every other open model on one-shot tasks - At pass@3: it did 75% -> 78.1%, matching GPT-5.6-Sol - Its precision remained stable, reporting fewer false positives than DeepSeek models The performance increase comes from a behavioral change: the new version is more persistent. It tends to run longer, and had a ~43% reasoning tokens increase. But the performance upgrade is worth that additional cost. 1/3 🧵

Holy moly: GLM-5.3 got much better in cybersecurity since our pre-release evaluation with Z.ai. It now matches GPT-5.6-Sol on our cybersecurity benchmark at 0.4x the cost 🤯 - At pass@1: it went from 60.4% to 65.6% CVEs rediscovered, crushing every other open model on one-shot tasks - At pass@3: it did 75% -> 78.1%, matching GPT-5.6-Sol - Its precision remained stable, reporting fewer false positives than DeepSeek models The performance increase comes from a behavioral change: the new version is more persistent. It tends to run longer, and had a ~43% reasoning tokens increase. But the performance upgrade is worth that additional cost. 1/3 🧵

34,641 просмотров

Videos

Больше нет контента для загрузки