正在加载视频...

视频加载失败

Introducing Devin Security Swarm A more cost effective and accurate way to find security vulnerabilities in complex codebases, based on a new architecture: Agentic MapReduce.

688,437 次观看 • 3 个月前 •via X (Twitter)

37 条评论

Cognition 的头像
Cognition3 个月前

In testing, Devin Security Swarm found 36 of 50 real-world GHSA vulnerabilities at 30% lower cost per finding than the next most accurate alternative.

Cognition 的头像
Cognition3 个月前

We built a new architecture for whole-codebase reasoning that we’re calling Agentic MapReduce. Security scanning is different from most coding tasks: a report is only trustworthy if the whole codebase is considered. But most agentic systems struggle to scale reasoning across large repos. Devin maps relevant signals across the repo, fans out focused agents over bounded shards, reduces their findings into one report, then verifies serious vulnerabilities in isolated sandboxes before marking them confirmed.

Cognition 的头像
Cognition3 个月前

The result is simultaneously more efficient and more accurate than other tools. We evaluated a variety of security scanning tools on a dataset of 50 GHSA vulnerabilities across 14 languages including Go, Rust, Python, Ruby, Java, C#, JavaScript, C, Swift, Dart, and Elixir. The dataset spans opens source repos of various sizes and of many software categories. Beyond excelling on our eval, Devin Security Swarm also found critical vulnerabilities that other tools missed, like a PHP sandbox bypass via template injection, an argument injection through metadata value parsing, and an overly broad deserialization surface.

Cognition 的头像
Cognition3 个月前

Security Swarm is a new pillar of Devin for Security: a suite of tools to help you find vulnerabilities, validate their exploitability at runtime, and ship remediation PRs. Learn more and try it today at:

Cognition 的头像
Cognition3 个月前

We’re also publishing extensive documentation and technical materials about Agentic MapReduce, including a deep-dive on our evals. Read our announcement: Learn about Agentic MapReduce: Check out the evals:

Alex Kaplan 的头像
Alex Kaplan3 个月前

@ido_pesok and Angela are goated Amazing launch!! Let's go

Gavin Baker 的头像
Gavin Baker3 个月前

Super cool

Michelle Liu 的头像
Michelle Liu3 个月前

the best team! 🤝 @ido_pesok @angelacareylin @(nick wong)

kylie chang 𓇢𓆸 的头像
kylie chang 𓇢𓆸3 个月前

the most GOATED TEAM!!

Gaspar Garcia 的头像
Gaspar Garcia3 个月前

I know that guy

Alex L 的头像
Alex L3 个月前

Nice work! System seems aligned with what I shared in my article for @RampLabs last month - have y'all had the chance to read it? Would love to share ideas and excited to see what Cog ships next 🙂

Wise 的头像
Wise3 个月前

Devin is cooking

Sarah Xu 的头像
Sarah Xu3 个月前

@RonMiasnik Dream team

Alberto Nunez-Garcia 的头像
Alberto Nunez-Garcia3 个月前

Very excited to try this after having used Codex security already, another pass should be useful! How do we use/invoke this?

Shoriful Dev 🔲 的头像
Shoriful Dev 🔲3 个月前

The interesting part isn't the vuln-finding, it's that every confirmed finding gets reproduced in an isolated sandbox before Devin opens the PR. That's the difference between a scanner that flags noise and one whose output you can actually trust to auto-merge.

Arnav Gupta 的头像
Arnav Gupta3 个月前

Awesome guys, but finding what's already in the codebase is one problem the other one is what the coding agent does while it's writing that code like pulling a poisoned package, running a shell command it shouldn't, acting on a prompt injection buried in a file @prismor_dev watches the agent's tool calls in real time and stops it there :)

Katie Cheng 的头像
Katie Cheng3 个月前

@stevenkplus1 the best! 🙌🙌 @ido_pesok @angelacareylin

Charity Quinn 的头像
Charity Quinn3 个月前

🙌🏻🙌🏻

Meryem Arik 的头像
Meryem Arik3 个月前

@devmchheda Background agents becoming mainstream

Evan Luke 的头像
Evan Luke3 个月前

Looks cool! you guys should share the list of vulns or the full benchmark at this point in time to contribute to open source and let others compare.

Raven 的头像
Raven3 个月前

agentic MapReduce is just 'we sent a lot of interns' with better branding

monto 的头像
monto3 个月前

IDO 🖤🐐

Evan Bacon 🥓 的头像
Evan Bacon 🥓3 个月前

🚀🚀🚀

AI Mastery Guide 的头像
AI Mastery Guide3 个月前

Agentic MapReduce for finding security vulnerabilities sounds like a smart way to scale code review across huge codebases.

Ofek Shaked | AI Engineer 的头像
Ofek Shaked | AI Engineer3 个月前

MapReduce helps bound the reasoning scope but the final verification sandbox still has to catch every false positive the fan out agents introduced. That last mile remains the expensive part in production.

Dünya Baradari 的头像
Dünya Baradari3 个月前

@ember_arlynx

Deniz Birlikci 的头像
Deniz Birlikci3 个月前

@Adhyyan security is becoming more and more important, Devin Security Swarm allows you to catch vulnerabilities at scale for less cost. one of the most important launches for the company imo

Saïd Aitmbarek 的头像
Saïd Aitmbarek3 个月前

dope concept guys

Hova 的头像
Hova3 个月前

Hey @grok kendi modellerimi

cv usk 的头像
cv usk3 个月前

Reproducing actual exploits in a sandbox to filter down to only real vulnerabilities feels genuinely production-grade. In security ops drowning in false positives, auto-opening the fix PR too could make backlog burndown dramatically faster.

Sajid ✘ 的头像
Sajid ✘3 个月前

congrats on this achievement

Trifon Getsov 的头像
Trifon Getsov3 个月前

agent swarms for code security is the right direction, scaling verification without humans in the loop is how this actually gets solved.

Tomas 的头像
Tomas3 个月前

Wow

Vijit Dhingra 的头像
Vijit Dhingra3 个月前

sick idea! We launched this to find bugs in applications with @getlark exactly 2 months before

Alan Wang 的头像
Alan Wang3 个月前

为什么未来我们绝对需要 100 倍以上的 AI 推理算力?AI 算力的真正吞噬者并不是人类在和 ChatGPT 聊天,而是正在席卷各行各业的“Agentic MapReduce”

Spencer Zhao 的头像
Spencer Zhao3 个月前

The bounded shard part is what makes this believable. Security review needs broad coverage, but every finding still has to collapse back to a reproducible exploit or failing check.

Inflectiv AI ⧉ 的头像
Inflectiv AI ⧉3 个月前

Applying a mapreduce framework to agentic workflows is a brilliant solution to context window limits in large repositories. This architectural change significantly lowers false positive rates, making automated scanning far more actionable for development teams.

相关视频