正在加载视频...
视频加载失败
3.7 Flash brings a big jump in agentic performance and coding accuracy. To demonstrate, we set up a 3-agent team to autonomously train a robotics control model from scratch. We hope you like 3.7 Flash, and you can read more here:
76,359 次观看 • 1 个月前 •via X (Twitter)
11 条评论

Today we're launching Gemini 3.7 Flash - our latest workhorse model for coding and agentic workflows, with an introductory price at half the original cost of 3.6 Flash. ⚡️ We have been iterating rapidly with the Flash series, going from 3.5 to 3.7 in just 3 months, making it more helpful across a wide range of tasks: • Software Engineering (DeepSWE v1.1): 37.0% ➔ 65.3% • Web Development (Code Arena Elo): 1506 ➔ 1588 • Enterprise Automation (AutomationBench): 13.4% ➔ 30.4%

how did you make sure it didn't just look up a repo on a similar version of that problem? There are a gazillion git repos with Mujoco control tasks with learned policies and parameters -- this doesn't make the model look particularly strong.

No one cares about your in-house controlled demos. Google is always creating heavy edited marketing content. The model speaks for itself by adoption and user reviews.

The 3-agent team is the interesting part. Once agents have distinct roles and a shared outcome, coordination becomes as important as capability. We’re going to learn a lot about team design from machines and vice versa.

Half the price with major performance gains 🚀 The automation benchmark jump from 13.4% to 30.4% is huge for Industry 4.0 applications.

The three-agent robotics demo is a much better stress test than another coding benchmark.

When are you going to fix it eventually looping quotation marks all over the place that degrade performance like hell on other tasks?

Genuinely wild watching that 3-agent robotics team work. We actually went deeper on the agentic capabilities here:

Interesting we got a ES optimizer!

Hi @koraykv , huge fan of your work at Google Labs! I have a concept for an AI interview preparation tool that cures interview anxiety. I’d love to share a quick summary with you or your team!

An autonomous 3-agent team training a robotics control model from scratch as the launch demo is the right choice. Flash tier usually gets toy demos; pairing it with multi-day agentic training is what proves the workhorse claim, not the price tag.




