Загрузка видео...

Не удалось загрузить видео

На главную

This week, we tested 3 latest models in our Game Arena Benchmark: → O3 → O4-mini → Gemini 2.5 Flash Across 4 games—Phoenix Wright, Sokoban, Candy Crush, and 2048—O3 dominated the zero-shot leaderboard, ranking #1 or #2 in nearly every task and outperforming previous SOTA models like O3-mini and...

15,105 просмотров • 1 год назад •via X (Twitter)

Комментарии: 9

Фото профиля Hao AI Lab
Hao AI Lab1 год назад

Leaderboard Snapshot: • O3 — 🥇 in Sokoban & 2048, 🥈 in Phoenix Wright & Candy Crush. • O4-mini — 🥇 in Candy Crush, lower ranks elsewhere. • Gemini 2.5 Flash — mid-tier across all games. O3 showed strong multi-modal reasoning, especially in spatial and long-horizon tasks, consistently placing in the top 2 across all games. O3 completely outperformed previous SOTA models like Gemini-2.5-pro and O3-mini. Gemini-2.5-flash delivered steady, cost-efficient performance without leading. For full results, check out the leaderboard here:

Фото профиля Hao AI Lab
Hao AI Lab1 год назад

OpenAI’s latest models—O3 and O4-mini—both showed strong reasoning in Candy Crush. They didn't just play; they optimized—aiming to eliminate the most candies per move. 🍬 This task demands more than quick reaction: → Vision to understand the board 👁️ → Reasoning to evaluate all possible swaps 🤔 → Planning to pick the best move 🎯 O3 and O4-mini consistently pushed beyond basic matches, prioritizing smart plays. O4-mini averaged 129, just shy of our human-level baseline of 134, and well above the previous SOTA O3-mini at 106.3. 👀

Фото профиля Hao AI Lab
Hao AI Lab1 год назад

We’re building more transparent, robust, and creative AI benchmarks—and we’d love your input. 💡📊 Got ideas for game-based evaluations? Have a challenge in mind? Drop your thoughts below! We also welcome more people to join our community, whether you're a researcher, builder, or just curious. → Leaderboard: → GitHub: → Website: Let’s push AI evaluation forward—together. 🚀

Фото профиля PowerBeatsVR
PowerBeatsVR3 лет назад

Get ready for a full-body VR workout that’s fun, fast, and intuitive — Play PowerBeatsVR (Now on Meta Quest) 🔥

Фото профиля Dylan Wolfe
Dylan Wolfe1 год назад

I don’t know why there isn’t any comments. This is awesome! It shows that o3 is making real world tangible progress! Please don’t stop doing these!!

Фото профиля Hao AI Lab
Hao AI Lab1 год назад

thank you!

Фото профиля Chase Brower
Chase Brower1 год назад

Great work! o3 looks like a great model.

Фото профиля Hao AI Lab
Hao AI Lab1 год назад

indeed, the visual reasoning reflects well on game play!

Фото профиля J Sam🌐
J Sam🌐1 год назад

Very good It should be tested to play more mobile app games >pubg , freefire >dragonball legends, dokkan battle >pokemon unite, go, masters >sports betting apps >stardew valley, farmville and sims

Похожие видео

We benchmarked leading multimodal foundation models (GPT-4o, Claude 3.5 Sonnet, Gemini, Llama, etc.) on standard computer vision tasks—from segmentation to surface normal estimation—using standard datasets like COCO and ImageNet. These models have made remarkable progress; however, it is unclear exactly where they stand in terms of understanding vision in detail. Especially when it comes to tasks beyond question-answering. How well do they understand an object's segments or geometry? Our analyses yield an assessment that is quantitatively and qualitatively detailed and is compatible with evaluations developed in the field of computer vision over the past decades. Observed trends: 🔹 The foundation models consistently underperform task-specific SOTA models across all tasks. However, they are respectable generalists, which is remarkable as they are presumably trained primarily on image-text-based tasks. 🔹 They perform semantic tasks notably better than geometric ones. 🔹 GPT-4o performs the best among non-reasoning models, getting the top position in 4 out of 6 tasks. 🔹 Reasoning models, e.g., o3, show improvements in geometric tasks. 🔹 The 'image generation' models, e.g., GPT-40 Image Generation, which have been natively trained multimodally, exhibit quirks. E.g., hallucinated objects, misalignment between the input and output, etc. 🔹 While the prompting techniques affect performance, better models exhibit less sensitivity to variations in prompts. We control for the variance introduced by the prompting methods in our experiments. 🌐 Detailed analyses, visualizations: ⌨️ code: 🧵 1/n

Amir Zamir

73,244 просмотров • 1 год назад

Which LLM reasons best when it doesn't have all the information? Enter LLM Poker Arena to find out. It's a Poker Playing benchmark where top reasoning models play Texas Hold'em poker against each other. Claude Opus 4.5, GPT-5.2, Gemini 2.5 Pro, and Grok 4 all sit at the same table and play full tournaments to see who finishes with the chips. Poker is very different when it comes to reasoning. It has to balance probabilistic reasoning, opponent modeling and make decisions under uncertainty. Poker is an interesting evaluation because it tests reasoning under incomplete information, something most coding benchmarks do not capture. In this tournaments the rules are: - Each LLM starts with $1,000 chips - Small and big blinds start at $25 / $50 - Blinds double every 3 minutes - All models run in their reasoning or thinking modes After the first 5 tournaments: - Claude Opus 4.5 with Thinking has 3 wins - GPT-5.2 has 2 wins - Grok 4 and Gemini 2.5 Pro have 0 wins Early results suggest Claude performs quite well at poker as well. Also five is a very small sample size. Planning to run many more tournaments, publish the full benchmark data and add a prediction market on top of it. Thanks for the suggestion clipz. Much more coming as part of Poker Cities !! This was built on Replit ⠕ using their AI integrations, which made it straightforward to connect Claude, GPT, and Gemini. What model do you think wins after 100 tournaments?

Anshul Dhawan

32,192 просмотров • 7 месяцев назад

GEMINI 3 LAUNCH IS HERE I got a SNEAK PEEK at Gemini 3 with Logan Kilpatrick (Google Deepmind), and it might be the most POWERFUL vibe-coding tool on the planet. A little breakdown: 1. Anyone can build 3D and casual games now You can vibecode full, playable 3D video games generated in minutes. Actual games with physics, characters, controls, and loops you can remix instantly. Pure insanity. I can see founders and brands spinning up games on the fly to ride trends and drive growth. 2. Intelligent apps are becoming the default We built apps where reasoning, memory, and multi-step planning were baked in from the start. Once you’re building apps with ACTUAL intelligence baked in, there’s a whole wave of new opportunities that weren’t possible before. 3. Gemini acts like a creative partner You describe the idea, Gemini fills in the gaps, challenges decisions, proposes alternatives, and iterates in real time. 4. Vibe coding hits a new level Gemini 3 can generate assets, code, game logic, UI, and narrative in one flow. Tools like Claude and Cursor feel fast. This feels like the next layer, the one where a single builder can compete with full teams. Logan Kilpatrick and I pushed Google Gemini 3 hard, and the outputs were solid. A few times we had to give it a few extra prompts but it took feedback really well. I think 1 year ago, a lot of people discounted Google in the AI arms race. Can you discount them anymore? Doubt it. After this, it feels like they at best leading, at worst leading. What do you think of Google's AI efforts/Gemini 3 My biggest takeaway was how it just felt like Gemini 3 had a little more vibe coding horsepower than anything I’ve used.

GREG ISENBERG

73,804 просмотров • 9 месяцев назад

Unstructured Thoughts about OpenAI o3, the nature of AGI, and Post-Labor Economics AGI just crossed a threshold—here’s why that matters and what we can do with it. I’ve been hammering on OpenAI’s new o3 model for a few days, long enough to watch the hype settle into something more interesting: utility. Benchmarks suggest a polite incremental bump; lived experience says we’ve entered a qualitatively different regime. o3 is the first model that feels faster than my ability to absorb its output. My brain—not the AI—has become the bottleneck. A new ceiling for human cognition? Most discussions of “alien intelligence” forget that we share the same sandbox: mathematics, physics, code, natural language. What shifts is cognitive horizon—the totality you can mentally represent and manipulate. o3 expands that horizon in real time. In an afternoon it consolidated two years of my work on post‑labor economics, stress‑tested the logic, surfaced data sources, and offered to autogenerate the Python notebooks. The cost of insight has collapsed from years to hours. If you merely outsource thought, you’ll stagnate. If you treat the model as a sparring partner—interrogating, refining, iterating—you’ll compound your own intelligence. Exponential leverage is now a choice, not a privilege. What o3 got right about my health project? I dumped the entire history of my chronic‑fatigue recovery protocol—including the five‑axis “burnout pentagram”—into memory and asked the model where I’d gone astray. It corrected a handful of minor assumptions and, more importantly, recalibrated my timeline: six‑to‑eight months of recovery left instead of eighteen. That’s not “replace your doctor” advice; it’s proof that large‑context reasoning is finally clinically useful. Post‑Labor Economics: the sketch that o3 and I built in one sitting 1. Metric 1 – Economic Agency Index (EAI) Income decomposed into wages, property, and transfers. The higher the property share, the more “post‑labor” you already are. 2. Metric 2 – Collective Purchasing Power (CPP) How much capital a county can mobilize without taxation or new debt. Rising CPP means you are compounding local prosperity. Interventions happen at the county level (subsidiarity): solar co‑ops in Arizona, riverfront greenways in the Midwest, data‑center dividends in fiber‑rich exurbs. Ownership is local, revenue is distributed, migration equilibrates naturally, and environmental stewardship becomes self‑interest rather than moral theater. UBI morphs from last‑ditch transfer to one of several levers for raising EAI. The bigger picture: AGI isn’t an oracle descending from the sky; it’s a time‑compression engine. Every minute you spend learning how to learn with it buys you an hour you would have burned doing rote synthesis. The frontier question is no longer “Will the machines replace us?” but “How fast can we upgrade ourselves in partnership with them?” What’s next? I’m cleaning the data, building the national EAI/CPP dashboard, and pressure‑testing the whole framework. I’ll publish the notebooks (or let o3 do it) once the numbers are solid. Meanwhile, I want to hear from you: Where does o3 add the most leverage in your world? Which of the post‑labor metrics feels wrong—or dangerously right? What failure mode should falsify this thesis? Drop your critique, your data source, or your wild counter‑proposal in the comments. Let’s map the edge of this new cognitive horizon together. —Dave

David Shapiro (L/0)

45,581 просмотров • 1 год назад

Today, I had the privilege of speaking to the HUMAIN team at our CEO Townhall. I often say this, and I mean it deeply: I am living my dream... working in a country where energy goes beyond oil. It’s in the people, the ambition, and the belief in building a better future. Saudi Arabia is unlike anywhere else. The hospitality is second to none. The vision is bold. And the commitment to shaping the future is real. What we are building at HUMAIN is foundational. We are not only participating in the AI era, but redefining it. Shifting the narrative from experimentation to real value creation. But more importantly, we are not afraid to challenge what an organization should look like in an AI-native world. In fact, we are pioneering it. We are actively reshaping how we work: - Moving toward a future where AI agents do the work - Empowering our people to focus on high-value thinking, creativity, and decision-making - Building strong foundations in security, governance, and guardrails - Continuously enhancing our models, systems, and operating frameworks This is not theory. This is happening now. What I saw today from our teams gives me absolute confidence: - Products that are not just innovative, but game-changing - Teams building at a speed and quality that challenge global norms - A culture focused on execution, ownership, and impact At #LEAP2026 this year, we will go beyond vision. We will show: Real demos. Real products. Some of them are a first of their kind in the world. And they are built right here in Saudi Arabia. I could not be more proud of this team. What they have accomplished in such a short time is remarkable. This is just the beginning. The future is not something we wait for. It’s something we build. #HUMAIN #LEAP #TheEndOfLimits #AI

Tareq Amin

17,856 просмотров • 4 месяцев назад

The frontier labs will crush you. That’s the first thing first friends say when I tell them I’m working on agents for biological research. It’s what I thought too before I started learning about biological agents. Now I’m convinced this space will be dominated by small teams building narrowly scoped, domain-specific models that uniquely embrace the complexity of living systems. So I’m diving in with a team of scientists and engineers who spent years watching models fail in real research settings - and built the “scar tissue” to understand why. We’re building Applied Scientific Intelligence, and some of the early results are encouraging: our models for biological data analysis & literature rank #1 on global benchmarks - outperforming ChatGPT, Claude, Gemini, as well as AI science companies with deep pockets (we’re still bootstrapped ;)). This recently caught the attention of NVIDIA and we partnered with them on our latest model: a new literature agent called “Alexandria” that’s going live in alpha this coming week. DM me if you want to try it out (or sign up in the link below). My main task now is building the distribution engine to get our agents in the hands of biologists. Estimates put the number of researchers at 9-10M globally. I think this number spikes in the next years as science starts to compound like software. Find an exponential curve and get in front of it, they say. Here goes.

Nate Hindman

13,890 просмотров • 4 месяцев назад

Why General AI Fails Tax Professionals — and What Makes TaxGPT State-of-the-Art Tax LLM Every tax professional using general-purpose AI tools like ChatGPT or Claude for research is one hallucinated source away from costly damage to their clients and professional career. So we ran a head-to-head benchmark: TaxGPT vs. OpenAI, Anthropic, and Gemini. Our methodology: 1️⃣ Quantitative testing on CPA + EA tax-specific exam questions 2️⃣ Qualitative testing complex, real-world scenarios across trust & estate, direct and indirect taxes, payroll, multi-state nexus, advisory, and compliance Results (Accuracy): TaxGPT: 96% Gemini 2.5 Pro: 89% Claude Sonnet 4: 87% GPT-4o: 80% GPT5: 92% But here’s the real story 👇 The Source Quality Gap General LLMs are great test takers. But tax research isn’t a trivia game — it’s a liability-driven profession where every position must be defensible. When we examined the sources behind their answers, the gap became a canyon: 🔍 OpenAI included citations only 9% of the time. 🔍 91% of answers had no verifiable source trail. 🔍 When citations did appear, they averaged 3.78 sources — many from Wikipedia, CNBC, AP News, Time, NerdWallet, etc. TaxGPT averages 14 authoritative sources per answer, all from primary/secondary tax law — IRC, Treasury Regs, court cases, IRS rulings, and our proprietary tax knowledge library. More details in our official blog post in the following post.

Kash from TaxGPT.com

59,253 просмотров • 9 месяцев назад

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 просмотров • 3 месяцев назад