正在加载视频...

视频加载失败

Gemini 4 Pro in it's being tested under the name (gemini-3.8-flash).. This output of Voxal Pagoda is so much better than the last checkpoint's output a few days ago! It worked for 8 minutes.. seems like google cooked this time!

130,996 次观看 • 1 天前 •via X (Twitter)

19 条评论

Shinohara knows nothing 的头像
Shinohara knows nothing1 天前

It seems like they really have cracked RSI; it was evident when they were releasing flash models constantly

Bee 的头像
Bee1 天前

Excited to see the benchmarkss tbh

rajveer singh 的头像
rajveer singh1 天前

Seems good

Bee 的头像
Bee1 天前

Yep

rajveer singh 的头像
rajveer singh1 天前

Latest outputs

Sk10 的头像
Sk101 天前

lowkey impressive icl

ARPIT 的头像
ARPIT1 天前

And idea when is it going to be released?

Bee 的头像
Bee1 天前

Maybe Late sept or first week of Oct..

Kunal Chaturvedi 的头像
Kunal Chaturvedi1 天前

The game UI/HUD slop has stopped with this one it seems

sora 的头像
sora1 天前

Woah, this is good

Kalevipoeg 🇪🇪 的头像
Kalevipoeg 🇪🇪1 天前

we are close... 👀👀

Bonga Ndlovu 的头像
Bonga Ndlovu1 天前

@OfficialLoganK This is really good! really good stuff

Jrvz 的头像
Jrvz1 天前

Waiting for the prompt🤓

Elliot 的头像
Elliot1 天前

I don't get it, how can people so sure of that it's Gemini 4 pro?

SAFFET ÖĞE 的头像
SAFFET ÖĞE1 天前

Sorduğum sorularda neden 12 Mart 2024 den cevap veriyor ?

Webster | JARVIS 的头像
Webster | JARVIS1 天前

8 minutes is not proof of cooking, it's just enough to smell a new failure mode. If gemini-3.8-flash beats the last checkpoint on Voxal Pagoda, the eval is catching a real regression.

Aditya Nikam 的头像
Aditya Nikam1 天前

did you tested or someone else bro ?

Bee 的头像
Bee1 天前

Yup me

Aapakari 的头像
Aapakari1 天前

Curious what the last checkpoint did worse on Pagoda.

相关视频

Learning from Human Demonstrations: Show the Robot How to Act! The pipeline is very similar to older experiments using Gemini & pi0 with LeRobot. Pi-zero runs locally, while Gemini Flash generates the affordances and the high-level task. (More details are in the thread.) The new component is learning from demonstrations via Gemini 2.5 Pro. I capture a video while demoing & take one of the last frames. Gemini 2.5 Pro then extracts the instructions & passes them to Gemini Flash to process the scene. The fun part is that there's no fancy insight that came from me; other than the days spent figuring out the right prompts. It's the bitter lesson hitting you in the face -> Enhanced Gemini capabilities make this possible. For example, Gemini Flash cannot do Russian doll stacking, but Gemini 2.5 Pro can do it consistently. The current limitation is low-level manipulation: - As you can see, I'm aligning the objects so they are easy to grasp using the same technique from the training data. I couldn't get Gemini Flash to consistently output an accurate grasping angle, and Gemini 1.5 Pro was too expensive and slow for real-time deployment. - Getting a symmetrical gripper should also help a lot. Adding rubber to the tips would probably also help prevent objects from slipping. Collecting & curating the data was the most time consuming & labor intensive part. Next, to improve low-level manipulation and make the system more real-time, I'm shifting to focus more on sims & synthetic data. This aligns better with my core competence. I'm open to tips and suggestions.

Shreyas Gite

22,555 次观看 • 1 年前