正在加载视频...

视频加载失败

This week, we tested 3 latest models in our Game Arena Benchmark: → O3 → O4-mini → Gemini 2.5 Flash Across 4 games—Phoenix Wright, Sokoban, Candy Crush, and 2048—O3 dominated the zero-shot leaderboard, ranking #1 or #2 in nearly every task and outperforming previous SOTA models like O3-mini and...

15,105 次观看 • 1 年前 •via X (Twitter)

9 条评论

Hao AI Lab 的头像
Hao AI Lab1 年前

Leaderboard Snapshot: • O3 — 🥇 in Sokoban & 2048, 🥈 in Phoenix Wright & Candy Crush. • O4-mini — 🥇 in Candy Crush, lower ranks elsewhere. • Gemini 2.5 Flash — mid-tier across all games. O3 showed strong multi-modal reasoning, especially in spatial and long-horizon tasks, consistently placing in the top 2 across all games. O3 completely outperformed previous SOTA models like Gemini-2.5-pro and O3-mini. Gemini-2.5-flash delivered steady, cost-efficient performance without leading. For full results, check out the leaderboard here:

Hao AI Lab 的头像
Hao AI Lab1 年前

OpenAI’s latest models—O3 and O4-mini—both showed strong reasoning in Candy Crush. They didn't just play; they optimized—aiming to eliminate the most candies per move. 🍬 This task demands more than quick reaction: → Vision to understand the board 👁️ → Reasoning to evaluate all possible swaps 🤔 → Planning to pick the best move 🎯 O3 and O4-mini consistently pushed beyond basic matches, prioritizing smart plays. O4-mini averaged 129, just shy of our human-level baseline of 134, and well above the previous SOTA O3-mini at 106.3. 👀

Hao AI Lab 的头像
Hao AI Lab1 年前

We’re building more transparent, robust, and creative AI benchmarks—and we’d love your input. 💡📊 Got ideas for game-based evaluations? Have a challenge in mind? Drop your thoughts below! We also welcome more people to join our community, whether you're a researcher, builder, or just curious. → Leaderboard: → GitHub: → Website: Let’s push AI evaluation forward—together. 🚀

PowerBeatsVR 的头像
PowerBeatsVR3 年前

Get ready for a full-body VR workout that’s fun, fast, and intuitive — Play PowerBeatsVR (Now on Meta Quest) 🔥

Dylan Wolfe 的头像
Dylan Wolfe1 年前

I don’t know why there isn’t any comments. This is awesome! It shows that o3 is making real world tangible progress! Please don’t stop doing these!!

Hao AI Lab 的头像
Hao AI Lab1 年前

thank you!

Chase Brower 的头像
Chase Brower1 年前

Great work! o3 looks like a great model.

Hao AI Lab 的头像
Hao AI Lab1 年前

indeed, the visual reasoning reflects well on game play!

J Sam🌐 的头像
J Sam🌐1 年前

Very good It should be tested to play more mobile app games >pubg , freefire >dragonball legends, dokkan battle >pokemon unite, go, masters >sports betting apps >stardew valley, farmville and sims

相关视频

We benchmarked leading multimodal foundation models (GPT-4o, Claude 3.5 Sonnet, Gemini, Llama, etc.) on standard computer vision tasks—from segmentation to surface normal estimation—using standard datasets like COCO and ImageNet. These models have made remarkable progress; however, it is unclear exactly where they stand in terms of understanding vision in detail. Especially when it comes to tasks beyond question-answering. How well do they understand an object's segments or geometry? Our analyses yield an assessment that is quantitatively and qualitatively detailed and is compatible with evaluations developed in the field of computer vision over the past decades. Observed trends: 🔹 The foundation models consistently underperform task-specific SOTA models across all tasks. However, they are respectable generalists, which is remarkable as they are presumably trained primarily on image-text-based tasks. 🔹 They perform semantic tasks notably better than geometric ones. 🔹 GPT-4o performs the best among non-reasoning models, getting the top position in 4 out of 6 tasks. 🔹 Reasoning models, e.g., o3, show improvements in geometric tasks. 🔹 The 'image generation' models, e.g., GPT-40 Image Generation, which have been natively trained multimodally, exhibit quirks. E.g., hallucinated objects, misalignment between the input and output, etc. 🔹 While the prompting techniques affect performance, better models exhibit less sensitivity to variations in prompts. We control for the variance introduced by the prompting methods in our experiments. 🌐 Detailed analyses, visualizations: ⌨️ code: 🧵 1/n

Amir Zamir

73,398 次观看 • 1 年前

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,272 次观看 • 1 个月前

GEMINI 3 LAUNCH IS HERE I got a SNEAK PEEK at Gemini 3 with Logan Kilpatrick (Google Deepmind), and it might be the most POWERFUL vibe-coding tool on the planet. A little breakdown: 1. Anyone can build 3D and casual games now You can vibecode full, playable 3D video games generated in minutes. Actual games with physics, characters, controls, and loops you can remix instantly. Pure insanity. I can see founders and brands spinning up games on the fly to ride trends and drive growth. 2. Intelligent apps are becoming the default We built apps where reasoning, memory, and multi-step planning were baked in from the start. Once you’re building apps with ACTUAL intelligence baked in, there’s a whole wave of new opportunities that weren’t possible before. 3. Gemini acts like a creative partner You describe the idea, Gemini fills in the gaps, challenges decisions, proposes alternatives, and iterates in real time. 4. Vibe coding hits a new level Gemini 3 can generate assets, code, game logic, UI, and narrative in one flow. Tools like Claude and Cursor feel fast. This feels like the next layer, the one where a single builder can compete with full teams. Logan Kilpatrick and I pushed Google Gemini 3 hard, and the outputs were solid. A few times we had to give it a few extra prompts but it took feedback really well. I think 1 year ago, a lot of people discounted Google in the AI arms race. Can you discount them anymore? Doubt it. After this, it feels like they at best leading, at worst leading. What do you think of Google's AI efforts/Gemini 3 My biggest takeaway was how it just felt like Gemini 3 had a little more vibe coding horsepower than anything I’ve used.

GREG ISENBERG

73,838 次观看 • 10 个月前

🚀Introducing VisualWebBench: A Comprehensive Benchmark for Multimodal Web Page Understanding and Grounding. 🤔What's this all about? Why this benchmark? > Back in Nov 2023, when we released MMMU ( a comprehensive multimodal understanding benchmark, we received feedback that it included very few UI screenshots. Considering the growing importance of UI understanding, especially with the rise of powerful agents like Devin ( which is built on the strong vision capability of #GPT4, we recognized the need for a benchmark focused on UI screenshot understanding.📸👀 > Multimodal #LLMs have significantly boosted web agents' performance on benchmarks like Mind2Web and WebArena. For instance, the SeeAct agent ( showcases the power of integrating vision into web agents. However, these benchmarks primarily evaluate the end-to-end task execution ability of web agents rather than their understanding of web pages. 🌉 Bridging the Gap with VisualWebBench > To provide a comprehensive evaluation of multimodal LLMs' web page understanding capabilities, we introduce VisualWebBench. Our benchmark spans 139 websites 🌐 across 12 domains 🏷️ and 87 sub-domains 🔍, ensuring a diverse and representative dataset. It assesses MLLMs at three levels: website-level, element-level, and action-level 📊, and encompasses seven tasks designed to evaluate understanding, OCR, grounding, and reasoning abilities 🧠💡. 😮 Surprising Findings > 🎉 Open-source models are catching up: Even though closed-source MLLMs are still leading the leaderboard, we are happy to see open-source models like LLaVA 1.6 34B achieve comparable performance to Gemini Pro. > 🧠 Grounding ability, crucial for developing MLLM-based web applications, is a weakness for most MLLMs. > 🖼️ Importance of Image Resolution: The limited image resolution handling capabilities of most open-source MLLMs restrict their utility in web scenarios, where rich text and elements are prevalent. > 🧱 Relatively strong correlation with general understanding benchmarks like MMMU but weak correlation with web agent benchmarks like Mind2Web. Web agent benchmarks primarily evaluate the end-to-end task execution ability of web agents, which involves a series of actions to accomplish a goal. In contrast, VisualWebBench emphasizes evaluating the foundational skills of MLLMs such as understanding and grounding web page elements. 💡Fun Fact > Claude Sonnet is better than Opus on our benchmark :) 🎓 Conclusion > VisualWebBench serves as a valuable resource for the community, driving research and development in the field of multimodal web page understanding and grounding. As MLLMs continue to evolve and improve, we look forward to seeing new applications and breakthroughs. We believe that our benchmark will contribute to the development of more powerful MLLMs in the web domain, ultimately leading to a more intuitive and efficient user experience on the web. Kudos to the student leads Junpeng Liu Yifan Song and the team Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li! 👏 Check out more details in the Junpeng's thread👇

Xiang Yue

56,696 次观看 • 2 年前

Unstructured Thoughts about OpenAI o3, the nature of AGI, and Post-Labor Economics AGI just crossed a threshold—here’s why that matters and what we can do with it. I’ve been hammering on OpenAI’s new o3 model for a few days, long enough to watch the hype settle into something more interesting: utility. Benchmarks suggest a polite incremental bump; lived experience says we’ve entered a qualitatively different regime. o3 is the first model that feels faster than my ability to absorb its output. My brain—not the AI—has become the bottleneck. A new ceiling for human cognition? Most discussions of “alien intelligence” forget that we share the same sandbox: mathematics, physics, code, natural language. What shifts is cognitive horizon—the totality you can mentally represent and manipulate. o3 expands that horizon in real time. In an afternoon it consolidated two years of my work on post‑labor economics, stress‑tested the logic, surfaced data sources, and offered to autogenerate the Python notebooks. The cost of insight has collapsed from years to hours. If you merely outsource thought, you’ll stagnate. If you treat the model as a sparring partner—interrogating, refining, iterating—you’ll compound your own intelligence. Exponential leverage is now a choice, not a privilege. What o3 got right about my health project? I dumped the entire history of my chronic‑fatigue recovery protocol—including the five‑axis “burnout pentagram”—into memory and asked the model where I’d gone astray. It corrected a handful of minor assumptions and, more importantly, recalibrated my timeline: six‑to‑eight months of recovery left instead of eighteen. That’s not “replace your doctor” advice; it’s proof that large‑context reasoning is finally clinically useful. Post‑Labor Economics: the sketch that o3 and I built in one sitting 1. Metric 1 – Economic Agency Index (EAI) Income decomposed into wages, property, and transfers. The higher the property share, the more “post‑labor” you already are. 2. Metric 2 – Collective Purchasing Power (CPP) How much capital a county can mobilize without taxation or new debt. Rising CPP means you are compounding local prosperity. Interventions happen at the county level (subsidiarity): solar co‑ops in Arizona, riverfront greenways in the Midwest, data‑center dividends in fiber‑rich exurbs. Ownership is local, revenue is distributed, migration equilibrates naturally, and environmental stewardship becomes self‑interest rather than moral theater. UBI morphs from last‑ditch transfer to one of several levers for raising EAI. The bigger picture: AGI isn’t an oracle descending from the sky; it’s a time‑compression engine. Every minute you spend learning how to learn with it buys you an hour you would have burned doing rote synthesis. The frontier question is no longer “Will the machines replace us?” but “How fast can we upgrade ourselves in partnership with them?” What’s next? I’m cleaning the data, building the national EAI/CPP dashboard, and pressure‑testing the whole framework. I’ll publish the notebooks (or let o3 do it) once the numbers are solid. Meanwhile, I want to hear from you: Where does o3 add the most leverage in your world? Which of the post‑labor metrics feels wrong—or dangerously right? What failure mode should falsify this thesis? Drop your critique, your data source, or your wild counter‑proposal in the comments. Let’s map the edge of this new cognitive horizon together. —Dave

David Shapiro (L/0)

45,581 次观看 • 1 年前

Today, I had the privilege of speaking to the HUMAIN team at our CEO Townhall. I often say this, and I mean it deeply: I am living my dream... working in a country where energy goes beyond oil. It’s in the people, the ambition, and the belief in building a better future. Saudi Arabia is unlike anywhere else. The hospitality is second to none. The vision is bold. And the commitment to shaping the future is real. What we are building at HUMAIN is foundational. We are not only participating in the AI era, but redefining it. Shifting the narrative from experimentation to real value creation. But more importantly, we are not afraid to challenge what an organization should look like in an AI-native world. In fact, we are pioneering it. We are actively reshaping how we work: - Moving toward a future where AI agents do the work - Empowering our people to focus on high-value thinking, creativity, and decision-making - Building strong foundations in security, governance, and guardrails - Continuously enhancing our models, systems, and operating frameworks This is not theory. This is happening now. What I saw today from our teams gives me absolute confidence: - Products that are not just innovative, but game-changing - Teams building at a speed and quality that challenge global norms - A culture focused on execution, ownership, and impact At #LEAP2026 this year, we will go beyond vision. We will show: Real demos. Real products. Some of them are a first of their kind in the world. And they are built right here in Saudi Arabia. I could not be more proud of this team. What they have accomplished in such a short time is remarkable. This is just the beginning. The future is not something we wait for. It’s something we build. #HUMAIN #LEAP #TheEndOfLimits #AI

Tareq Amin

17,856 次观看 • 5 个月前