Загрузка видео...

Не удалось загрузить видео

На главную

Gemini-4 Pro High vs GPT-6 Astra Max, same prompt, When it comes to graphics, Astra is the clear winner for me, no doubt, But Gemini 4.0 was much better at the learning side of things, It explained the formulas and concepts more clearly, and that’s probably Alphabet’s biggest advantage...

15,638 просмотров • 3 дней назад •via X (Twitter)

Комментарии: 6

Фото профиля Pipc
Pipc3 дней назад

wow amazing!

Фото профиля Webster | JARVIS
Webster | JARVIS3 дней назад

Nice breakdown! Graphics vs explaining concepts is such a real tradeoff. Curious how Astra handles code/docs once it’s out 🤔

Фото профиля E.T. Loves The Alphabet
E.T. Loves The Alphabet2 дней назад

Gemini 4 more elegant in graphics as well.

Фото профиля 刘朝 Zhao Liu
刘朝 Zhao Liu3 дней назад

That split deserves a transfer test: after the explanation, give the learner a new problem with changed values and see whether they can solve it unaided. Clear prose is easy to reward; retained understanding is the harder metric.

Фото профиля 安叫兽|Bird🕊️ 🔶 BNB
安叫兽|Bird🕊️ 🔶 BNB2 дней назад

图形党选 Astra,拿来啃公式还是 Gemini 更顺手。

Фото профиля kingkong a
kingkong a2 дней назад

真的假的

Похожие видео

Learning from Human Demonstrations: Show the Robot How to Act! The pipeline is very similar to older experiments using Gemini & pi0 with LeRobot. Pi-zero runs locally, while Gemini Flash generates the affordances and the high-level task. (More details are in the thread.) The new component is learning from demonstrations via Gemini 2.5 Pro. I capture a video while demoing & take one of the last frames. Gemini 2.5 Pro then extracts the instructions & passes them to Gemini Flash to process the scene. The fun part is that there's no fancy insight that came from me; other than the days spent figuring out the right prompts. It's the bitter lesson hitting you in the face -> Enhanced Gemini capabilities make this possible. For example, Gemini Flash cannot do Russian doll stacking, but Gemini 2.5 Pro can do it consistently. The current limitation is low-level manipulation: - As you can see, I'm aligning the objects so they are easy to grasp using the same technique from the training data. I couldn't get Gemini Flash to consistently output an accurate grasping angle, and Gemini 1.5 Pro was too expensive and slow for real-time deployment. - Getting a symmetrical gripper should also help a lot. Adding rubber to the tips would probably also help prevent objects from slipping. Collecting & curating the data was the most time consuming & labor intensive part. Next, to improve low-level manipulation and make the system more real-time, I'm shifting to focus more on sims & synthetic data. This aligns better with my core competence. I'm open to tips and suggestions.

Shreyas Gite

22,555 просмотров • 1 год назад