Загрузка видео...

Не удалось загрузить видео

На главную

3.7 Flash brings a big jump in agentic performance and coding accuracy. To demonstrate, we set up a 3-agent team to autonomously train a robotics control model from scratch. We hope you like 3.7 Flash, and you can read more here:

76,359 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 11

Фото профиля koray kavukcuoglu
koray kavukcuoglu1 месяц назад

Today we're launching Gemini 3.7 Flash - our latest workhorse model for coding and agentic workflows, with an introductory price at half the original cost of 3.6 Flash. ⚡️ We have been iterating rapidly with the Flash series, going from 3.5 to 3.7 in just 3 months, making it more helpful across a wide range of tasks: • Software Engineering (DeepSWE v1.1): 37.0% ➔ 65.3% • Web Development (Code Arena Elo): 1506 ➔ 1588 • Enterprise Automation (AutomationBench): 13.4% ➔ 30.4%

Фото профиля POWNS
POWNS1 месяц назад

how did you make sure it didn't just look up a repo on a similar version of that problem? There are a gazillion git repos with Mujoco control tasks with learned policies and parameters -- this doesn't make the model look particularly strong.

Фото профиля media deleite
media deleite1 месяц назад

No one cares about your in-house controlled demos. Google is always creating heavy edited marketing content. The model speaks for itself by adoption and user reviews.

Фото профиля Jeff James Martin
Jeff James Martin1 месяц назад

The 3-agent team is the interesting part. Once agents have distinct roles and a shared outcome, coordination becomes as important as capability. We’re going to learn a lot about team design from machines and vice versa.

Фото профиля Felix Manufaktur
Felix Manufaktur1 месяц назад

Half the price with major performance gains 🚀 The automation benchmark jump from 13.4% to 30.4% is huge for Industry 4.0 applications.

Фото профиля Leo Lu
Leo Lu1 месяц назад

The three-agent robotics demo is a much better stress test than another coding benchmark.

Фото профиля Shepherdess
Shepherdess1 месяц назад

When are you going to fix it eventually looping quotation marks all over the place that degrade performance like hell on other tasks?

Фото профиля Michael Waitze
Michael Waitze1 месяц назад

Genuinely wild watching that 3-agent robotics team work. We actually went deeper on the agentic capabilities here:

Фото профиля Yinggan Xu
Yinggan Xu1 месяц назад

Interesting we got a ES optimizer!

Фото профиля Chirag Chaudhary
Chirag Chaudhary1 месяц назад

Hi @koraykv , huge fan of your work at Google Labs! I have a concept for an AI interview preparation tool that cures interview anxiety. I’d love to share a quick summary with you or your team!

Фото профиля Loong🐉
Loong🐉1 месяц назад

An autonomous 3-agent team training a robotics control model from scratch as the launch demo is the right choice. Flash tier usually gets toy demos; pairing it with multi-day agentic training is what proves the workhorse claim, not the price tag.

Похожие видео

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,360 просмотров • 1 месяц назад