正在加载视频...
视频加载失败
👾Watch a multi-agent system built with Gemini 3.6 Flash iterate on playable game design in real time. It uses Gemini 3.5 Flash-Lite to design, construct, and verify puzzle challenges based on live player actions balancing speed and reasoning in a fast-paced environment.
30,630 次观看 • 1 个月前 •via X (Twitter)
25 条评论

Developer guide:

I’m sorry. I’ve been using the 3.6 flash model in the few past days and he is very fast, that’s for sure. But he fails in simple tasks. Give him a simple job of programming a web page with some graphics and mostly he gets it all wrong.

The visible design → construct → verify loop is the interesting part. Agent systems earn trust when verification is a first-class stage. I’d love to see what the verifier rejects—and how disagreements or failed checks are surfaced instead of silently retried.

how do you handle the handoff when the lite model fails verification

finally a system that designs puzzles as slowly as i solve them

The verification step is the interesting part here. In multi-agent systems, coordination is often easier than knowing which agent produced a result worth trusting. How are conflicts between the designer, builder, and verifier resolved?

In your software development work, are you using Opus and Sol models to write this system as Google? I assume you did the other simple operations using Flash.

ai designing puzzles in real time so it can beat you and gaslight you about it

Gain Z

3.6 Flash? you running a beta or just typo'd the model name

i can already feel the AI outsmarting my next move

Worst phone making company google. No service no support. If you want this use it as secondary device. Mainly for indian coustmer. I am waiting from last five days for screen replacement.

Google, I would like to make a request: please do not let Gemini be driven by such haste for quick results—rushing things brings no benefit.

Are you excited for AI designed microtransactions?

The problem isn't Gemini Flash's speed... but its issue is announcing the completion of tasks that haven't necessarily been tested and proven. Others are a bit more measured in their responses.

Yes but WHAT HARNESS

ok but which agent actually gets the final say when the lite model disagrees with the flash on a puzzle

实时生成关卡这块,延迟表现挺想看

This is insane. Imagine dynamic difficulty scaling in games where the AI actually builds new levels on the fly. 😀

I thought it was an update about Jules

🎙️OPINION🎙️: NEW GEMINI, THE SAME PROBLEMS?⚠️ @Google @GoogleDeepMind @googleaidevs announces three new models while thousands of users continue to report, without response, problems that have persisted since previous versions. What you should know: ❌ Since the implementation of Gemini 3.5 Flash, users have documented increasingly vague and less detailed responses after each update across all Gemini models. No one at Google has officially acknowledged this. 📉 Releasing three models at once isn't always a sign of leadership: it can be a sign that none of them are quite ready. ⚠️ OpenAI, Anthropic, and Meta release improvements week after week. This pressure is evident in every rushed announcement from Google. 👉 The cybersecurity model only benefits governments. Ordinary users, who also need protection, are left out without explanation. 💡 Why it matters: When a company prioritizes advertising over fixing bugs, it's the average user who ends up paying the price.

flash-lite verifying flash - aren't you worried about the hallucination chain?

wait so the lite model is doing the design and the flash is just running it?

The model only learned how to play the game because it was given a narrow set of possible moves and clear win conditions. Real world multi agent systems need to handle ambiguity and incomplete information way better than this demo does.

..

