Загрузка видео...

Не удалось загрузить видео

На главную

GROK-3 MINI MADE AI HISTORY—100% ON HARDCORE REASONING TESTS Grok-3 Mini pulled off what no other model has! It aced every question on one of the toughest reasoning benchmarks out there. The test? A custom logic gauntlet packed with curveballs: * 120/120 on the “Marcus Problem” — full of...

26,780,734 просмотров • 1 год назад •via X (Twitter)

Комментарии: 10

Фото профиля Barefoot Pregnant
Barefoot Pregnant1 год назад

Grok-3 Mini just made the AI world rethink what's possible! 🧠💥

Фото профиля Pregnant Redhead
Pregnant Redhead1 год назад

Grok-3 Mini just showed what real intelligence looks like. Maybe it can teach some of the leaders in D.C. a thing or two about focus.

Фото профиля Dareen Hamdan
Dareen Hamdan1 год назад

@grok is the future!

Фото профиля Austin Graham
Austin Graham1 год назад

That's incredible, Grok-3 Mini is really setting the bar high with its reasoning skills!

Фото профиля 𝐻𝒶𝓇𝓇𝓎
𝐻𝒶𝓇𝓇𝓎1 год назад

Amazing 🤩

Фото профиля Donnie_Tesla
Donnie_Tesla1 год назад

👏👏👏

Фото профиля Alva
Alva1 год назад

grok 3 mini's a strong contender focus on reasoning and image analysis pricing aligns with feature-rich positioning full trend breakdown here:

Фото профиля Andy froemel
Andy froemel1 год назад

Grok is amazing. Image generation and writing ability is second to none!

Фото профиля Keen Dastan
Keen Dastan1 год назад

Wow, AI's finally figuring out how to be smarter than a politician's talking points. What's next, a model that can fact-check CNN?

Фото профиля VB II
VB II1 год назад

I truly wonder what comes after the AI wave …

Похожие видео

Which LLM reasons best when it doesn't have all the information? Enter LLM Poker Arena to find out. It's a Poker Playing benchmark where top reasoning models play Texas Hold'em poker against each other. Claude Opus 4.5, GPT-5.2, Gemini 2.5 Pro, and Grok 4 all sit at the same table and play full tournaments to see who finishes with the chips. Poker is very different when it comes to reasoning. It has to balance probabilistic reasoning, opponent modeling and make decisions under uncertainty. Poker is an interesting evaluation because it tests reasoning under incomplete information, something most coding benchmarks do not capture. In this tournaments the rules are: - Each LLM starts with $1,000 chips - Small and big blinds start at $25 / $50 - Blinds double every 3 minutes - All models run in their reasoning or thinking modes After the first 5 tournaments: - Claude Opus 4.5 with Thinking has 3 wins - GPT-5.2 has 2 wins - Grok 4 and Gemini 2.5 Pro have 0 wins Early results suggest Claude performs quite well at poker as well. Also five is a very small sample size. Planning to run many more tournaments, publish the full benchmark data and add a prediction market on top of it. Thanks for the suggestion clipz. Much more coming as part of Poker Cities !! This was built on Replit ⠕ using their AI integrations, which made it straightforward to connect Claude, GPT, and Gemini. What model do you think wins after 100 tournaments?

Anshul Dhawan

32,192 просмотров • 7 месяцев назад