Загрузка видео...

Не удалось загрузить видео

На главную

Gemini 3 Flash now uses an agentic "think-act-observe" loop to solve complex visual tasks 🤖 Google DeepMind engineer Paul Ruiz demonstrates how the model runs Python code automatically to zoom and inspect items, annotate images, and re-visualize data into charts.

106,963 просмотров • 7 месяцев назад •via X (Twitter)

Комментарии: 35

Фото профиля Google AI Developers
Google AI Developers7 месяцев назад

@ptruiz_dev Check out the documentation: Try it in

Фото профиля Luc Biggs
Luc Biggs7 месяцев назад

@GoogleDeepMind @ptruiz_dev AgenticVision .com

Фото профиля Gregor
Gregor7 месяцев назад

I'm starting to see this 'think-act-observe' loop as a way to automate the iterative design process, where the model refines its actions based on the results of its previous decisions, much like how a designer would refine their design based on user feedback. This could unlock new possibilities for interactive visualizations and real-time data exploration. Building something like this for a mobile finance app has given me a similar idea - I'm excited to explore this further.

Фото профиля Larry Panozzo
Larry Panozzo7 месяцев назад

@GoogleDeepMind @ptruiz_dev It has been fun to see Gemini 3 Flash do things no other model can do. That zoom & inspect loop is fantastic.

Фото профиля Ilpo Leppänen
Ilpo Leppänen7 месяцев назад

I absolutely love Gemini 3 Flash! What makes it so fun model to play with is that it's a true wildcard 🃏 I've often found that it navigates its way through the information space in various unexpected ways and comes up with great novel ideas, while still ripping through the task with a speedy punch! And sometimes you find it almost impossible to pursue a goal with strict instructions as if it faces even a slightest roadblock, it will turn into a frenzied Sonic, bouncing around to get through the obstacle. It's clearly a model made to iterate - fast

Фото профиля Raf
Raf7 месяцев назад

@GoogleDeepMind @ptruiz_dev Still can't find the cat 😺

Фото профиля Goblin
Goblin7 месяцев назад

@GoogleDeepMind @ptruiz_dev This just straight out of the box solved problems for me I previously could not solve with LLMs, even with complex flows. Insane work google, thanks! @ptruiz_dev

Фото профиля CrazyAI Tech
CrazyAI Tech7 месяцев назад

@GoogleDeepMind @ptruiz_dev Okay, an AI can zoom. Please make my chicken be Einstein.

Фото профиля ⚡Jessy SEO formation #IA générative #chatGPT 🦊
⚡Jessy SEO formation #IA générative #chatGPT 🦊7 месяцев назад

@GoogleDeepMind @ptruiz_dev @grok fais un tableau qui explique quel utilisation chaque modele de gemini est specialisé est bon

Фото профиля Professor. Dr. Ife. Indra Aria Dika E.bt. S.me.
Professor. Dr. Ife. Indra Aria Dika E.bt. S.me.7 месяцев назад

@GoogleDeepMind @ptruiz_dev Thank you

Фото профиля xiaoyuqin
xiaoyuqin7 месяцев назад

@GoogleDeepMind @ptruiz_dev @grok 总结视频的内容f 总结视频的内容 总结视频的内容a w

Фото профиля ameno
ameno7 месяцев назад

@GoogleDeepMind @ptruiz_dev I would try Gemini if you could create a damn api key. GCP/Ai studio, no key works despite enabling the correct services.

Фото профиля Ihsan
Ihsan7 месяцев назад

@GoogleDeepMind @ptruiz_dev I'm watching it so many questions have

Фото профиля Solve SF
Solve SF7 месяцев назад

@GoogleDeepMind @ptruiz_dev Does this work on @openrouter ?

Фото профиля 10turtle
10turtle7 месяцев назад

@GoogleDeepMind @ptruiz_dev Gemini 3 Flash using a “think → act → observe” loop is a big shift. 🤖 Running Python to zoom, inspect, annotate, and even turn visuals into charts automatically? That’s not just vision that’s reasoning + execution combined. AI isn’t just seeing anymore… it’s doing

Фото профиля mr.tweeps | IDN
mr.tweeps | IDN7 месяцев назад

@GoogleDeepMind @ptruiz_dev Impressive! Gemini 3 Flash’s ability to autonomously analyze and visualize complex data is fascinating.

Фото профиля Kerris Kahdelka
Kerris Kahdelka7 месяцев назад

@GoogleDeepMind @ptruiz_dev When can gemini in the browser have a "laser pointer" so it can point and draw on the web UI with us while we work. Would be pretty great to have gemini live + laser pointer almost like actual screensharing over zoom with Gemini.

Фото профиля bullbearlovechild
bullbearlovechild7 месяцев назад

@GoogleDeepMind @ptruiz_dev This sounds great but the demo in your video fails. The resolution is bad, but it's still clear that most pictures have the wrong label: Duck -> Zebra, Pig -> Giraffe, Hippo -> Bear, Crocodile -> Crab

Фото профиля Professor. Dr. Ife. Indra Aria Dika E.bt. S.me.
Professor. Dr. Ife. Indra Aria Dika E.bt. S.me.7 месяцев назад

@GoogleDeepMind @ptruiz_dev Great revolution!

Фото профиля BLOND BOND
BLOND BOND7 месяцев назад

@GoogleDeepMind @ptruiz_dev INSTANTLY THOUGHT WHEN SAW LAST EXAMPLE WAS MED TECH USE FOR POTENTIAL CANCEROUS MOLES... ANALYSIS AND GIVE LIKLEYHOOD OF CANCEROUS... MANY USE CASES... THATS WHY AI GOLD RUSH CAN BE FOR EVERYONE...

Фото профиля Muhammad Yusuf
Muhammad Yusuf7 месяцев назад

Gemini 3 Flash's think-act-observe loop is powered by native tool-use + on-device Python execution (via MediaPipe & LiteRT), allowing sub-second reasoning→code→vision cycles without external API calls. In the circuit board demo, it not only zooms and labels but can also run edge-detection filters or count components in real time. Early internal benchmarks show ~4.2× faster multimodal reasoning than Gemini 1.5 Flash on vision-heavy agent tasks. Huge step toward truly agentic on-device multimodal AI. Docs + cookbook examples already live in AI Studio — highly recommend trying the "visual data extraction" templates first.

Фото профиля Stew
Stew7 месяцев назад

@GoogleDeepMind @ptruiz_dev When it is out on paid API?

Фото профиля Charles |🟪$ETH🌟👑
Charles |🟪$ETH🌟👑7 месяцев назад

@GoogleDeepMind @ptruiz_dev @AIxC_Official @aixcfoundation #Tenki

Фото профиля john
john7 месяцев назад

@GoogleDeepMind @ptruiz_dev ChatGPT has been doing this for me for quite a while.

Фото профиля Dari Rinch
Dari Rinch7 месяцев назад

@GoogleDeepMind @ptruiz_dev Love this direction — moving from static vision to agentic loops is key. One missing piece is a deterministic COMMIT/NOCOMMIT step before executing Python/tool calls, with both outcomes logged for traceability. Spec + minimal implementation:

Фото профиля thirdweb
thirdweb7 месяцев назад

@GoogleDeepMind @ptruiz_dev now give it a wallet.

Фото профиля #INTERNETofAGENTS
#INTERNETofAGENTS7 месяцев назад

@GoogleDeepMind @ptruiz_dev 🫣

Фото профиля Uncle J
Uncle J7 месяцев назад

@GoogleDeepMind @ptruiz_dev “think-act-observe” is the new normal. The gap isn’t models anymore. It’s orchestration.

Фото профиля BuildrLab | AI & Cloud Solutions
BuildrLab | AI & Cloud Solutions7 месяцев назад

@GoogleDeepMind @ptruiz_dev This loop is the real unlock for vision: don’t “one-shot caption” — iterate with tools. Would love to see a run-log / trace standard here (actions, crops, code executed) so teams can debug + eval these agents like we do with software.

Фото профиля LitoManuelRuiz
LitoManuelRuiz7 месяцев назад

@GoogleDeepMind @ptruiz_dev Grok gemini cause IA PRAY

Фото профиля Arthur
Arthur7 месяцев назад

@GoogleDeepMind @ptruiz_dev when you guys are planning on releasing a stable version of flash, not preview ?

Фото профиля Hamad
Hamad7 месяцев назад

@GoogleDeepMind @ptruiz_dev @grok it's all computer vision?

Фото профиля Mateus Carrijo
Mateus Carrijo7 месяцев назад

@GoogleDeepMind @ptruiz_dev google is king!!! fantastic

Фото профиля mrkelly
mrkelly7 месяцев назад

@GoogleDeepMind @ptruiz_dev The real bottleneck isn't the model. It's knowing what to ask it to see.

Фото профиля Jedi⚡️
Jedi⚡️7 месяцев назад

@GoogleDeepMind @ptruiz_dev Gemini 3 Flash's new agentic loop turns vision into an active, code-driven investigation.

Похожие видео