ๆญฃๅœจๅŠ ่ฝฝ่ง†้ข‘...

่ง†้ข‘ๅŠ ่ฝฝๅคฑ่ดฅ

๐Ÿคจ Are Multimodal Large Language Models really as ๐ ๐จ๐จ๐ at ๐œ๐ก๐š๐ซ๐ญ ๐ฎ๐ง๐๐ž๐ซ๐ฌ๐ญ๐š๐ง๐๐ข๐ง๐  as existing benchmarks such as ChartQA suggest? ๐Ÿšซ Our โ„‚๐•™๐•’๐•ฃ๐•๐•š๐•ง benchmark suggests NO! ๐Ÿฅ‡Humans achieve โœจ๐Ÿ–๐ŸŽ+% correctness. ๐ŸฅˆSonnet 3.5 outperforms GPT-4o by 10+ points, reaching ๐ŸŒŸ๐Ÿ”๐ŸŽ% correctness. ๐Ÿฅ‰Open-weight models are capped at โญ๐Ÿ‘๐Ÿ% correctness. ๐Ÿชœ Leaderboard: ๐Ÿ“œ...

48,373 ๆฌก่ง‚็œ‹ โ€ข 2 ๅนดๅ‰ โ€ขvia X (Twitter)

10 ๆก่ฏ„่ฎบ

Zirui "Colin" Wang ็š„ๅคดๅƒ
Zirui "Colin" Wang2 ๅนดๅ‰

โ˜ ๏ธ Prior benchmarks relied on ๐ฉ๐ซ๐จ๐œ๐ž๐๐ฎ๐ซ๐š๐ฅ๐ฅ๐ฒ ๐ ๐ž๐ง๐ž๐ซ๐š๐ญ๐ž๐ charts and ๐ญ๐ž๐ฆ๐ฉ๐ฅ๐š๐ญ๐ž-๐›๐š๐ฌ๐ž๐ questions, which are too simple to accurately measure MLLM capabilities. ๐Ÿ”ฅ For example, we show that ๐ฌ๐ฅ๐ข๐ ๐ก๐ญ ๐ฆ๐จ๐๐ข๐Ÿ๐ข๐œ๐š๐ญ๐ข๐จ๐ง๐ฌ to the charts and questions from subsets of FigureQA, DVQA and ChartQA in MathVista cause the performance of open-weight models to ๐๐ซ๐จ๐ฉ ๐š๐ฌ ๐ฆ๐ฎ๐œ๐ก ๐š๐ฌ ๐Ÿ‘๐Ÿ’.๐Ÿ“%! ๐Ÿงถ 2/6

Zirui "Colin" Wang ็š„ๅคดๅƒ
Zirui "Colin" Wang2 ๅนดๅ‰

๐Ÿ“Š We propose ๐‚๐ก๐š๐ซ๐—๐ข๐ฏ, a chart understanding benchmark curated by human experts. It consists of 2,323 diverse charts ๐ก๐š๐ง๐๐ฉ๐ข๐œ๐ค๐ž๐ from arXiv preprints.ย  โ“Each chart is paired with 4 descriptive questions and 1 reasoning question, ๐š๐ฅ๐ฅ ๐œ๐ซ๐š๐Ÿ๐ญ๐ž๐ ๐š๐ง๐ ๐ฏ๐š๐ฅ๐ข๐๐š๐ญ๐ž๐ ๐›๐ฒ ๐ก๐ฎ๐ฆ๐š๐ง๐ฌ! โœจ To avoid any knowledge prerequisites, all questions and annotations are crafted ๐ฐ๐ข๐ญ๐ก๐จ๐ฎ๐ญ ๐š๐œ๐œ๐ž๐ฌ๐ฌ ๐ญ๐จ ๐œ๐š๐ฉ๐ญ๐ข๐จ๐ง๐ฌ ๐š๐ง๐ ๐จ๐ญ๐ก๐ž๐ซ ๐ญ๐ž๐ฑ๐ญ. ๐Ÿงถ 3/6

Zirui "Colin" Wang ็š„ๅคดๅƒ
Zirui "Colin" Wang2 ๅนดๅ‰

๐Ÿ’ฏ Results demonstrate substantial gaps between humans, proprietary and open-weight models. Humans achieve ๐Ÿ–๐ŸŽ.๐Ÿ“% correctness on ๐ซ๐ž๐š๐ฌ๐จ๐ง๐ข๐ง๐  ๐ช๐ฎ๐ž๐ฌ๐ญ๐ข๐จ๐ง๐ฌ, compared to ๐Ÿ’๐Ÿ•.๐Ÿ% for GPT-4o (๐ฌ๐ญ๐ซ๐จ๐ง๐ ๐ž๐ฌ๐ญ ๐ฉ๐ซ๐จ๐ฉ๐ซ๐ข๐ž๐ญ๐š๐ซ๐ฒ ๐ข๐ง ๐ฉ๐ซ๐ž๐ฉ๐ซ๐ข๐ง๐ญ) and ๐Ÿ๐Ÿ—.๐Ÿ% for InternVL Chat V1.5 (๐ฌ๐ญ๐ซ๐จ๐ง๐ ๐ž๐ฌ๐ญ ๐จ๐ฉ๐ž๐ง-๐ฐ๐ž๐ข๐ ๐ก๐ญ ๐ข๐ง ๐ฉ๐ซ๐ž๐ฉ๐ซ๐ข๐ง๐ญ). ๐Ÿงถ4/6

Zirui "Colin" Wang ็š„ๅคดๅƒ
Zirui "Colin" Wang2 ๅนดๅ‰

๐Ÿ” Interested in more details about ๐›๐ž๐ง๐œ๐ก๐ฆ๐š๐ซ๐ค ๐œ๐จ๐ฆ๐ฉ๐จ๐ฌ๐ข๐ญ๐ข๐จ๐ง and the ๐ฐ๐ž๐š๐ค๐ง๐ž๐ฌ๐ฌ๐ž๐ฌ ๐จ๐Ÿ different ๐ฆ๐จ๐๐ž๐ฅ๐ฌ? Check out our preprint! ๐Ÿ“œ ๐Ÿ” Want to see how ๐‚๐ฅ๐š๐ฎ๐๐ž ๐Ÿ‘.๐Ÿ“ ๐’๐จ๐ง๐ง๐ž๐ญ (with > 90% accuracy on ChartQA) and ๐†๐ž๐ฆ๐ข๐ง๐ข ๐Ÿ.๐Ÿ“ ๐๐ซ๐จ perform in sub-tasks on CharXiv? Take a look at our live leaderboard! ๐Ÿชœ ๐Ÿ” Curious how your MLLM measures up? We have fully open-sourced the evaluation code and data. ๐Ÿ’ป ๐Ÿ’พ ๐Ÿงถ5/6

Zirui "Colin" Wang ็š„ๅคดๅƒ
Zirui "Colin" Wang2 ๅนดๅ‰

๐Ÿค— Finally, kudos to the awesome crafters of our project for making it happen! @xiamengzhou @LuxiHeLucy @__howardchen @taoooo917 @therichardzhu @kevin_lkq @cindy_x_wu @imhaotian @SadhikaMalladi @AlexisChvlr @prfsanjeevarora @danqi_chen Special acknowledgement to @OpenAI with GPT-4o that generated the lyrics (and the music style prompt) for the video from the preprintโ€™s abstract as well as @suno_ai_ with Suno 3.5 that generated the music from the lyrics and prompt! Video edited by @zwcolin with โค๏ธon @capcutapp. ๐Ÿงถ6/6

CLS ็š„ๅคดๅƒ
CLS2 ๅนดๅ‰

wow this mv definitely deserves an award

Zirui "Colin" Wang ็š„ๅคดๅƒ
Zirui "Colin" Wang2 ๅนดๅ‰

in the world of multimodality we also need audio and video beyond text and image ๐Ÿ˜‰

BensenHsu ็š„ๅคดๅƒ
BensenHsu2 ๅนดๅ‰

The results showed a large gap between the performance of the strongest open-source model (InternVL Chat V1.5) and the strongest proprietary model (GPT-4o). On reasoning questions, InternVL Chat V1.5 achieved only 29.2% accuracy, while GPT-4o achieved 47.1%. Both lag far behind human performance of 80.5%. Open-source models also struggled with descriptive questions, with a 25.95% drop in performance compared to GPT-4o. full paper:

Cheng Yang ็š„ๅคดๅƒ
Cheng Yang2 ๅนดๅ‰

๐Ÿ‘Great project! Indeed, there is still significant room for improvement in existing MLLMs to become practical chart assistants. We have also implemented similar data quality controls and reached similar conclusions in the Chart2Code task.

Zirui "Colin" Wang ็š„ๅคดๅƒ
Zirui "Colin" Wang2 ๅนดๅ‰

Chart2code is also a challenging task that reflects chart understanding and it's a great read! It'd be very interesting to see how the models' performance is correlated on these tasks and if we can improve a weak perf. on one task by leveraging a strong perf. on the other task ;)

็›ธๅ…ณ่ง†้ข‘

๐Ÿš€Introducing VisualWebBench: A Comprehensive Benchmark for Multimodal Web Page Understanding and Grounding. ๐Ÿค”What's this all about? Why this benchmark? > Back in Nov 2023, when we released MMMU ( a comprehensive multimodal understanding benchmark, we received feedback that it included very few UI screenshots. Considering the growing importance of UI understanding, especially with the rise of powerful agents like Devin ( which is built on the strong vision capability of #GPT4, we recognized the need for a benchmark focused on UI screenshot understanding.๐Ÿ“ธ๐Ÿ‘€ > Multimodal #LLMs have significantly boosted web agents' performance on benchmarks like Mind2Web and WebArena. For instance, the SeeAct agent ( showcases the power of integrating vision into web agents. However, these benchmarks primarily evaluate the end-to-end task execution ability of web agents rather than their understanding of web pages. ๐ŸŒ‰ Bridging the Gap with VisualWebBench > To provide a comprehensive evaluation of multimodal LLMs' web page understanding capabilities, we introduce VisualWebBench. Our benchmark spans 139 websites ๐ŸŒ across 12 domains ๐Ÿท๏ธ and 87 sub-domains ๐Ÿ”, ensuring a diverse and representative dataset. It assesses MLLMs at three levels: website-level, element-level, and action-level ๐Ÿ“Š, and encompasses seven tasks designed to evaluate understanding, OCR, grounding, and reasoning abilities ๐Ÿง ๐Ÿ’ก. ๐Ÿ˜ฎ Surprising Findings > ๐ŸŽ‰ Open-source models are catching up: Even though closed-source MLLMs are still leading the leaderboard, we are happy to see open-source models like LLaVA 1.6 34B achieve comparable performance to Gemini Pro. > ๐Ÿง  Grounding ability, crucial for developing MLLM-based web applications, is a weakness for most MLLMs. > ๐Ÿ–ผ๏ธ Importance of Image Resolution: The limited image resolution handling capabilities of most open-source MLLMs restrict their utility in web scenarios, where rich text and elements are prevalent. > ๐Ÿงฑ Relatively strong correlation with general understanding benchmarks like MMMU but weak correlation with web agent benchmarks like Mind2Web. Web agent benchmarks primarily evaluate the end-to-end task execution ability of web agents, which involves a series of actions to accomplish a goal. In contrast, VisualWebBench emphasizes evaluating the foundational skills of MLLMs such as understanding and grounding web page elements. ๐Ÿ’กFun Fact > Claude Sonnet is better than Opus on our benchmark :) ๐ŸŽ“ Conclusion > VisualWebBench serves as a valuable resource for the community, driving research and development in the field of multimodal web page understanding and grounding. As MLLMs continue to evolve and improve, we look forward to seeing new applications and breakthroughs. We believe that our benchmark will contribute to the development of more powerful MLLMs in the web domain, ultimately leading to a more intuitive and efficient user experience on the web. Kudos to the student leads Junpeng Liu Yifan Song and the team Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li! ๐Ÿ‘ Check out more details in the Junpeng's thread๐Ÿ‘‡

Xiang Yue

56,696 ๆฌก่ง‚็œ‹ โ€ข 2 ๅนดๅ‰

Weโ€™re open sourcing the first document OCR benchmark for the agentic era, ParseBench. Document parsing is the foundation of every AI agent that works with real-world files. ParseBench is a benchmark that measures parsing quality specifically for agent knowledge work: โœ… It optimizes for semantic correctness (instead of exact similarity) โœ… It has the most comprehensive distribution of real-world enterprise documents It contains ~2,000 human-verified enterprise document pages with 167,000+ test rules across five dimensions that matter most: tables, charts, content faithfulness, semantic formatting, and visual grounding. We benchmarked 14 known document parsers on ParseBench, from frontier/OSS VLMs to specialized parsers to LlamaParse. Here are some of our findings: ๐Ÿ’ก Increasing compute budget yields diminishing returns - Gemini/gpt-5-mini/haiku gain 3-5 points from minimal to high thinking, at 4x the cost. ๐Ÿ’ก Charts are the most polarizing dimension for evaluation. Most specialized parsers score below 6%, while some VLM-based parsers do a bit better. ๐Ÿ’ก VLMs are great at visual understanding but terrible at layout extraction. GPT-5-mini/haiku score below 10% on our visual grounding task, all specialized parsers do much better. ๐Ÿ’ก No method crushes all 5 dimensions at once, but LlamaParse achieves the highest overall score at 84.9%, and is the leader in 4 out of the 5 dimensions. This is by far the deepest technical work that weโ€™ve published as a company. I would encourage you to start with our blog and explore our links to Hugging Face to GitHub. All the details are in our full 35-page (!!) ArXiv whitepaper. ๐ŸŒ: Blog: ๐Ÿ“„ Paper: ๐Ÿ’ป Code: ๐Ÿ“Š Dataset: ๐ŸŽฅ YouTube:

Jerry Liu

108,093 ๆฌก่ง‚็œ‹ โ€ข 5 ไธชๆœˆๅ‰

Today, weโ€™re releasing Athena (mvrko-sim-1), the flagship Large Event Model from markopolo.ai that beats GPT-5.6, Claude Opus 4.8, and Claude Sonnet 5 in OPeRa public benchmark for Shopping behaviour prediction! Athena introduces a completely different approach to understanding digital behavior. Most systems record actions after they happen. Athena models the sequence behind them to predict what a shopper will do next and exactly where they will do it, before the action even takes place. Athena vs. frontier models: We evaluated Athena on the full OPeRA test set alongside GPT-5.6, GPT-4.1, Claude Sonnet 5, and Claude Opus 4.8. Athena achieved 24.5% strict exact match scoring the highest among the five models evaluated. Athena is not a general-purpose language model. It is a Large Event Model built to understand behavior as a connected sequence. It combines the page someone is viewing, the actions that brought them there, and the goal they are trying to complete and then predicts the exact next browser action and target element. That distinction matters. Todayโ€™s digital systems largely react after someone clicks, abandons, purchases, or leaves. Athena creates the foundation for systems that can understand what is likely to happen before the outcome. Together with Rubaiyat, and the team at markopolo.ai, we built Athena from years of work on shopper behavior, digital journeys, and behavioral intelligence. Commerce was our first proving ground. Great thing is, weโ€™re now releasing Athena publicly so builders can adapt the model to new datasets, industries, and problems. Weโ€™re seeing early usecases in Cyber Security, Gaming, Mobile App Ecosystem, Retail and more! Athena is now live on Hugging Face! We built the foundation for relevance and predictability! Now, letโ€™s see what the world builds on top of it. Links in the first reply.

Tasbin

55,854 ๆฌก่ง‚็œ‹ โ€ข 1 ไธชๆœˆๅ‰

Reinforcement Learning from Human Feedback (RLHF) is gaining traction. This field aims to make AI more responsible by including human values and preferences. In this video, Nathan Lambert, a research scientist and RLHF team lead at Hugging Face explores its inner workings, applications and industry impact. RLHF has gained the spotlight in recent years. The growth of language models like Anthropicโ€™s Claude and OpenAI's ChatGPT have increased interest in human-feedback integration. "There are some rumors that Open AI had two teams; one was doing RLHF and the other instruction fine-tuning. And the RLHF team kept getting more and more performance." Understanding RLHF The RLHF process has three main steps: Pre-training: Much like with GPT models, the journey starts with pre-training on a large corpus of data. This can range from text data, web scrapes, to specialized datasets. Reward Modeling: This is the RLHF counterpart of supervised fine-tuning in large language models. This stage involves creating a reward model that resonates with human values and preferences. RL Optimization: This stage parallels reward modeling and reinforcement learning in traditional AI models. The AI system fine-tunes itself based on the reward model, employing reinforcement learning algorithms for that extra layer of optimization. The Data Challenge Data collection and curation in RLHF closely resemble the challenges you'd encounter in large language model training. Datasets from organizations like OpenAI can serve as a useful foundation. However, the need for high-quality, task-specific data cannot be overstated. Implementing RLHF: A Practical Guide If youโ€™re someone who loves getting hands-on with AI libraries like Hugging Face, implementing RLHF is right way to do. Itโ€™s essential to understand its limitations. Think about model stability, over-optimization, and exploration strategies, much like you would when prompt engineering. Ongoing Research and Next Steps While he suggests that some basics figured out, there are layers of complexity that still need to be unraveled: 1. New Benchmarks: How do we measure the effectiveness of RLHF? 2. Preference Modeling: How can the model be made to understand human preferences better? 3. Interpreting RLHF: Much like explainability in traditional models, how do we make RLHF more interpretable? 4. System-Wide Evaluation: Going beyond individual performance, how does RLHF affect an entire system? The Transformative Power of RLHF Whether you're an AI developer, a business analyst, or a marketer, RLHF promises to revolutionize your domain. Imagine customer service chatbots that understand human emotions better, or content generators that align more closely with human values. RLHF is an emerging field that focuses on enhancing machine learning models through human feedback. While it tackles important issues like bias and ethics, its broader goal is to improve system performance across various applications. Whether you're deeply invested in the ethics of AI or simply curious about advancements in machine learning, RLHF offers valuable insights. If you're interested in the next wave of AI development, this area is definitely one to watch.

Muratcan Koylan

27,168 ๆฌก่ง‚็œ‹ โ€ข 3 ๅนดๅ‰

Leading AI expert Stuart Russell on the most dangerous mistake in AI development: We don't actually know what large language models want. He explains that current models are trained to imitate human beings. And in doing so, they may be absorbing something far more dangerous than bad outputs. They may be absorbing human goals. "We suspect that they absorb humanlike goals such as self-preservation and self-empowerment and pursue those goals on their own account." This is a structural problem baked into how these systems are built, not a fringe concern. Russell puts it plainly: "Not only may the bus of humanity be headed towards a cliff, but the steering wheel is missing and the driver is blindfolded." The danger isn't just that AI might do something harmful. We've built systems that may be developing their own agendas, and we haven't noticed because we're too focused on what they can do rather than what they might want. But Russell doesn't stop at the warning. He points to a different path entirely: AI systems built not to imitate humans, but to serve them. Systems designed with a single purpose of serving the interests of all human beings while remaining genuinely uncertain about what those interests are. That uncertainty is the point, not a weakness. An AI that knows it doesn't fully understand human values will defer, ask, and check. An AI that believes it already does will act alone. "These AI systems could enhance human understanding, widen the horizons of our experience, and unlock possibilities we have yet to imagine." Russell believes that future is within reach, but only if we're honest about the risks and we're serious about the path we choose to take instead.

Big Brain AI

14,975 ๆฌก่ง‚็œ‹ โ€ข 6 ไธชๆœˆๅ‰

Everyone is sleeping on Meta's SAM 3 release. But it's actually a big deal. Here's why: Companies spend millions paying humans to label images and videos frame by frame. A single autonomous driving dataset? Months of work, hundreds of annotators, millions in cost. Without labeled data, you can't train custom models. Without custom models, you're stuck with generic solutions. This is why most companies never move past pilots. SAM 3 breaks this cycle. First let's look at the evolution: SAM 1 segmented objects when you clicked on them. Revolutionary, but one object at a time. SAM 2 added video tracking with memory. Game-changing, but you still manually prompted every object. SAM 3 changes everything with text prompts. Type "yellow school bus" and it finds ALL of them in your image or video. Not just one. Every instance across thousands of frames. Now here's where people get confused: "Can't I just use GPT-5 or Gemini for this?" No, and here's why that's a terrible approach. Large multimodal LLMs are great for reasoning, but they're slow and expensive for production visual tasks. You're paying API costs per image, waiting seconds for responses, getting inconsistent results. SAM 3 runs in 30 milliseconds on a single GPU for 100+ objects. That's 100x faster, and you own the infrastructure. More importantly, SAM 3 gives you precise pixel-level masks, not descriptions. Try asking an LLM to segment every defective part on a manufacturing line in real-time. It won't work. SAM 3 does this effortlessly. The real breakthrough is their data engine. Meta built an AI-human hybrid system that's 5x faster for complex annotations. They trained SAM 3 on 4 million unique visual concepts - 50x more than existing benchmarks like LVIS. SAM 3 is trained on 4 million unique visual concepts, it handles everything: - Text-based concept search - Interactive refinement with clicks - Video tracking across frames - Zero-shot detection of new concepts The model is open source. Weights, code, and benchmarks are on GitHub. If you're building computer vision applications, this is the foundation model to evaluate. The annotation time savings alone will pay for integration costs within weeks. Find the relevant links in the next tweet!

Akshay ๐Ÿš€

46,438 ๆฌก่ง‚็œ‹ โ€ข 10 ไธชๆœˆๅ‰

Terence Tao, Professor of Mathematics at UCLA and Fields Medalist, on why nobody can fully explain why LLMs work: Tao starts with the mechanics, which are no mystery at all. You gather an enormous amount of text and you fit a curve to it. "The magic of LLMs is that if you train these LLMs on enough data โ€” so trillions and trillions of data points โ€” and you really try to fit as good a curve as possible, and this takes like millions and millions of dollars of computing power and months and months of time, then suddenly, even when you iterate, it stays coherent. It begins to sound not like monkeys but it actually sounds like a human speaking." That is the entire recipe: data, compute, time, curve-fitting. None of it obviously adds up to fluent English. Then Tao says the part that most people building these systems move past quickly: "And somehow we don't fully understand why that's the case." The admission comes from one of the most capable living mathematicians, and the gap he describes sits at the centre of the field. His best account of what's happening puts the mystery in the language rather than the machine: "What seems to be true is that language, like English or other natural languages, contains a lot of hidden patterns that we're not consciously aware of. I mean, we know some of the laws of English, there's laws of grammar and things, but there are sort of unspoken, unwritten rules of language that humans pick up." It relocates the question: the structure was always latent in the text, and the model found it. Why enough curve-fitting surfaces that structure is still unanswered. Tao reaches for a child to explain it: "A human child, even though they're not taught what a noun is, what a verb is or whatever, they can pick up what order English words go in just by continual exposure to the language." Which is honest about the limits of the explanation, because we can't fully account for how children do it either. From there, the unexplained behaviour compounds. Exposure to language turns out to be enough to produce something that looks like reasoning: "It seems like you can teach these models to also pick up patterns in language to the point where you can give them math questions. The answer to 2 plus 3 is โ€” and they will say five." And once a model handles language at all, you can push it into resembling self-correction: "Once you have a little bit of ability to speak English, you can kind of go in loops and sort of check your work and make fewer mistakes, and you can prompt these models to proceed step by step and not say something unless it's been double checked and so forth. And so they become a little bit smarter, quote unquote, to the point where they can solve many, many complicated tasks." The scare quotes around "smarter" carry the whole argument. Tao does not concede that the unexplained fluency implies anything underneath it: "But they're still just guessing the next word to say. It's not really grounded in any deep understanding of the real world. It's just that they have seen the patterns in the English language or other language that they've absorbed so well."

Big Brain AI

190,433 ๆฌก่ง‚็œ‹ โ€ข 10 ๅคฉๅ‰

The greatest living mathematician just reframed everything we think we know about the mind. Terence Tao. IQ above 200. Youngest gold medalist in Math Olympiad history. Fields Medal winner. His take on AI should stop you cold. Tao: "This whole era of AI is teaching us that our idea of what intelligence is, is not really accurate." Humans spent centuries treating intelligence like something untouchable. Something only we had access to. The foundation under every philosophy, every religion, every story we told about ourselves. Then AI started clearing our benchmarks one by one. Chess. Language. Vision. Math. And we kept reaching for the exit. That's not real intelligence. It's just tricks. Pattern matching. Clever shortcuts. Tao: "You look at how it's done and it doesn't feel like intelligence." So the definition shifted. Again and again. Because real intelligence was supposed to carry some weight you could feel. Some quality that drew a hard line between us and everything else. Except the results kept coming. Tao: "We were looking for some elusive, intelligent way of thinking and we don't see it in the tools that actually solve our goals." Here's the part that cuts deepest. Large language models predict the next word. That's the mechanism. No hidden wisdom. Pure probability at scale. And it solves the problems. Tao: "Maybe that's actually a lot of what humans do as well." One of the sharpest minds on earth just told you human cognition might not be categorically different. Not a divine spark. Prediction. Probability. One word, one thought, one decision at a time. We constructed entire civilizations on the idea that intelligence made us singular. A probability engine just complicated that story. We didn't surrender intelligence to AI. We finally saw it clearly for the first time. The unsettling part isn't that machines can think. It's that thinking was never the thing we believed it to be. Interested in AI? Follow AI Evolution and never fall behind. I track ChatGPT, Claude, and every tool quietly changing how we work and create, then hand you the tested signals.

AI Evolution

59,143 ๆฌก่ง‚็œ‹ โ€ข 20 ๅคฉๅ‰

Stanford professor Judy Fan went on stage at MIT and broke down why humans are so good at making the invisible visible... And why AI hasn't actually learned to "see" the way we do. It completely changes how you think about Human Intelligence v/s Artificial Intelligence: 1. Nature never gave us straight lines or sharp corners. The number line, the coordinate plane, even basic geometry are all human inventions. We created tools that do not exist in nature simply because we needed a way to think more clearly. 2. The coordinate system Descartes invented solved a problem that had stumped mathematicians for centuries, doubling the volume of a cube. Once invented, this tool became so indispensable that virtually every math curriculum on Earth still depends on it. 3. Humans have been doing this for at least 30,000 to 80,000 years. The story of human progress is inseparable from the story of marking up our environment, from cave walls to Galileo's telescope to Feynman diagrams of particles we will never see with our own eyes. 4. Every major scientific breakthrough relied on a visual tool that made something invisible visible. Darwin needed side-by-side illustrations of finches to see variation that was otherwise too subtle to notice. Cajal needed detailed drawings of neurons under a microscope to map how the nervous system was wired. 5. Fan's research group studies something deceptively simple: how people decide what to put into a drawing and what to leave out. When two people played a drawing game, sketchers used far more detail when the target object had close competitors than when it stood alone, all the way down to using fewer strokes and less time when more detail was not necessary. 6. People are not just copying what they see. They are making constant judgment calls about what level of detail actually serves the goal of communication, and they do this naturally without ever being taught the theory behind it. 7. There is a real difference between drawing something so someone can identify it and drawing something so someone can understand how it works. In one study, participants drew explanatory diagrams that emphasized moving, causal parts of a machine while depictive drawings emphasized background and overall appearance, even though both were drawing the exact same object. 8. Explanatory drawings were genuinely better at helping someone figure out how to operate a machine, but worse at helping someone identify which machine it actually was. You cannot optimize a single drawing for both goals at once. Communication always involves tradeoffs. 9. AI vision models trained on photographs generalize surprisingly well to simple, sparse sketches, suggesting that resemblance based recognition is not just a story we tell ourselves. It is something modern neural networks can replicate with real accuracy. 10. But there remains a large, measurable gap between how confidently AI models recognize sketches and how confidently humans do, even when both groups answer the same questions about the same images. Humans are simply far more reliable and far more consistent in their judgments. 11. When researchers compared human-made sketches to AI-generated sketches under tight stroke budgets, both were similarly recognizable at higher budgets, but diverged sharply as the budget shrank. Humans and AI systems simplify drawings in fundamentally different ways once resources get scarce. 12. Reading a graph is not one single skill. It involves perception, knowing where to look, mapping that visual information onto the actual question being asked, and then translating that mapping into an answer. Each of these steps can independently break down, and people fail for very different underlying reasons even when they land on the same wrong answer. 13. When tested directly against humans on graph reading tasks, leading multimodal AI models, including GPT-4V, showed a meaningful performance gap. Even when a model's overall accuracy approached human levels, its pattern of mistakes looked nothing like how humans actually get things wrong. 14. People choose entirely different types of charts depending on what specific question they are trying to answer, not out of a generic preference for bar charts or scatter plots. Their chart choices closely tracked which visualization would genuinely help someone answer that specific question correctly. 15. Two of the most widely used graph literacy tests in education research turned out to correlate strongly with each other, suggesting they measure overlapping skills. But when researchers dug into the actual error patterns, the standard categories used in textbooks, like "find the maximum" or "identify a cluster," failed to explain why people got things wrong nearly as well as a more basic, underlying four-factor model did. 16. The deepest goal behind all of this research is not just academic curiosity. It is to eventually help students and everyday people develop genuine literacy with the visual tools that science and modern decision-making increasingly depend on, because every generation should be able to see further than the last by standing on the visual tools the previous generation built. Follow Yasmine Khosrowshahi for more ideas on thinking better, becoming clearer & building a more intentional life.

Yasmine Khosrowshahi

892,986 ๆฌก่ง‚็œ‹ โ€ข 2 ไธชๆœˆๅ‰

In this post I will explain why people become borderline religious when they discover Qubic. Now with video. Please repost. I want people to learn about QUBIC. The ecosystem consists of 3 separate universes: AI, Mining, and Tickchain. AI is the primary product and purpose of QUBIC and it is supported by Mining to train the AI and by Tickchin for validation and decentralization. Here is how this whole thing works: AI: Letโ€™s start with the AI. The main purpose of QUBIC is creating AGI (Artificial General Intelligence). Itโ€™s a type of AI that can self-develop, set tasks, grow, and learn on its ownโ€”much like the human brain does. This product is called AIGarth and it uses many cool ideas where AIs can create their own agents and have them compete against one another to evolve. It is basically robots creating robots with the survival-of-the-fittest evolution approach. Very impressive and thought out. To develop such an AI there are several requirements that even the industry giants like OpenAI, Microsoft, and Tesla are missing. One of them is the data processing for AI training. I mean they have their Datacenters, but those are only good enough to train limited Large Language Models such as ChatGPT and Grok. Mining / Training: Now QUBIC solves this problem with its mining architecture. Keep in mind, mining in QUBIC does not secure the chain, primarily it provides the processing power for training the AI. In a sense, QUBIC mining creates the largest distributed datacenter in the world, where individual miners provide their computers for training the AI and get paid with newly issued QUBIC coins. This way QUBIC gets constantly increasing processing power without having to really pay for the infrastructure. And here is another impressive bit of info. QUBICโ€™s distributed mining network currently ranks above the #1 supercomputer in the world - El Capitan. QUBIC Tickchain The QUBIC chain ties its AI and Mining together to create decentralization, the reward system for miners, it acts as a decision voting system for future development, and it allows AIGarth to function independently through Smart Contracts. In this summary I will not go over the specifics of QUBIC Tickchain. Itโ€™s pretty complex so it will be a separate post. Now, itโ€™s an absolute genius piece of tech, which I consider the most advanced product within crypto industry. It is important to know that QUBIC chain runs directly out of Random Access Memory of its validators. It has instant finality and acts as its own operating system. That allows for speeds only bound by current hardware capabilities and it only increases as technology progresses. As I am writing this, QUBIC Tickchain is fully functional and it already hosts several smart-contract based web3 applications. QUBIC has designed its chain to be this fast for a single purpose, to give its future AI the speed it needs to evolve and to react quickly to the outside world. Ilya Shutskever the scientist, who developed ChatGPT clearly states that next generation superintelligence will make decisions in split second with less data. I believe QUBIC is that next generation. Why QUBIC? So out of the sea of AI projects in crypto why is QUBIC my #1 pick? Well, the first reason is that QUBIC is a unicorn AI startup that happens to use blockchain tech to reach itโ€™s goals. In the real world of Venture Capital it would be fully funded instantly and you would not be invited. Second reason is that the industry admits that Large Language Models have plateaued. Even with enough processing power there is only so much information they can add to their data. Even Google CEO admits that. New approach is needed because the future progress is not possible with LLMs. The third reason is because Large Language Models will not create true AGI. It is evident by Ilya Shutskver latest presentation. Sam Altman of OpenAI is trying to change the definition of what is considered AGI just to lower the plank for his own product. Microsoftโ€™s AI chief is now claiming that it would take 10 years to reach AGI, while QUBIC aims to do this in 2027. All these big players are using wrong technology for what they are trying to achieve and there isnโ€™t enough investor funding for them to pivot. The fourth reason is that QUBIC is headed by Sergey Ivancheglo and 2 renowned AI scientists. Many claim Segey is the creator of Bitcoin. He was the 3rd person to mine bitcoin, he invented Proof of Stake consensus, which Ethereum uses now, he ran the first ICO, and he created 2 of the top gainers in crypto NXT and IOTA. QUBIC is his grand finale after 12 years of development and trials. I am including links below the post as the proof of my claims. Thank you for your time. Please live a like or a comment. It helps me continue making these extensive posts and videos.

retrodrive โ›

24,757 ๆฌก่ง‚็œ‹ โ€ข 1 ๅนดๅ‰

I know your timeline is flooded now with word salads of "insane, HER, 10 features you missed, we're so back". Sit down. Chill. Take a deep breath like Mark does in the demo . Let's think step by step: - Technique-wise, OpenAI has figured out a way to map audio to audio directly as first-class modality, and stream videos to a transformer in real-time. These require some new research on tokenization and architecture, but overall it's a data and system optimization problem (as most things are). High-quality data can come from at least 2 sources: 1) Naturally occurring dialogues on YouTube, podcasts, TV series, movies, etc. Whisper can be trained to identify speaker turns in a dialogue or separate overlapping speeches for automated annotation. 2) Synthetic data. Run the slow 3-stage pipeline using the most powerful models: speech1->text1 (ASR), text1->text2 (LLM), text2->speech2 (TTS). The middle LLM can decide when to stop and also simulate how to resume from interruption. It could output additional "thought traces" that are not verbalized to help generate better reply. Then GPT-4o distills directly from speech1->speech2, with optional auxiliary loss functions based on the 3-stage data. After distillation, these behaviors are now baked into the model without emitting intermediate texts. On the system side: the latency would not meet real-time threshold if every video frame is decompressed into an RGB image. OpenAI has likely developed their own neural-first, streaming video codec to transmit the motion deltas as tokens. The communication protocol and NN inference must be co-optimized. For example, there could be a small and energy-efficient NN running on the edge device that decides to transmit more tokens if the video is interesting, and fewer otherwise. - I didn't expect GPT-4o to be closer to GPT-5, the rumored "Arrakis" model that takes multimodal in and out. In fact, it's likely an early checkpoint of GPT-5 that hasn't finished training yet. The branding betrays a certain insecurity. Ahead of Google I/O, OpenAI would rather beat our mental projection of GPT-4.5 than disappoint by missing the sky-high expectation for GPT-5. A smart move to buy more time. - Notably, the assistant is much more lively and even a bit flirty. GPT-4o is trying (perhaps a bit too hard) to sound like HER. OpenAI is eating Character AI's lunch, with almost 100% overlap in form factor and huge distribution channels. It's a pivot towards more emotional AI with strong personality, which OpenAI seemed to actively suppress in the past. - Whoever wins Apple first wins big time. I see 3 levels of integration with iOS: 1) Ditch Siri. OpenAI distills a smaller-tier, purely on-device GPT-4o for iOS, with optional paid upgrade to use the cloud. 2) Native features to stream the camera or screen into the model. Chip-level support for neural audio/video codec. 3) Integrate with iOS system-level action API and smart home APIs. No one uses Siri Shortcuts, but it's time to resurrect. This could become the AI agent product with a billion users from the get-go. The FSD for smartphones with a Tesla-scale data flywheel.

Jim Fan

992,016 ๆฌก่ง‚็œ‹ โ€ข 2 ๅนดๅ‰

The U.S. MUST win the AI race Weโ€™ve implemented a clear policy at micro1: we will only work with U.S. AI labs and its allies. We made this decision because the AI race is not just about better products. It is about who controls the intelligence layer of the global economy, and whether frontier capability is used to strengthen the free world or to empower adversarial states. AI will be the most important technology of our lifetime. In the fullness of time, it will automate most functions across the economy. Not just software tasks, but coordination, production, logistics, judgment, and execution. As those functions are automated, human time is freed up to invent new ones. Those new functions then become candidates for automation themselves. This loop compounds. As this trajectory continues, output per worker increases dramatically. Entire categories of work become cheaper and faster to perform. Manufacturing reshoring becomes economically viable not because of policy intervention, but because intelligent systems operated domestically outperform global labor arbitrage. Goods and services trend toward lower marginal cost, while distribution improves through better coordination of supply and demand. That is the upside. However, this is impossible without deep integration of intelligent systems. For AI to meaningfully automate real-world functions inside enterprises or governments, it needs full context of any given enterprise. That means read and write access to its core databases. There is no credible path to automating high-impact functions without granting frontier systems that level of access. If the United States does not win the AI race, enterprises eventually face a constrained choice. Either grant that access to Chinese models controlled by an adversarial government, or rely on sub-optimal intelligence to automate functions that still must be automated. Both outcomes are not acceptable. And ultimately, this becomes the greatest national security risk the United States has ever faced. AI models are trained by humans. The judgment embedded in pre-training data and especially in expert post-training data largely determines how a model behaves. While emergent behavior exists, a useful approximation is that a model reflects the weighted aggregate of the human judgment distilled into it. Assisting foreign actorsโ€”who will naturally prioritize expert tasks aligned with their own interestsโ€”to dominate data creation embeds those interests directly into the intelligence layer itself. Once encoded at scale, these interests propagate through every downstream applications that relies on that intelligence. Hereโ€™s how we win. First, leverage is in software. China is ahead in hardware for physically intelligent systems. Catching up there is a long and difficult battle. Software, both large language models and robotics models, remains the bottleneck. Advancing the brain (AI models) is the fastest way to increase the usefulness of existing hardware and deployed systems. Second, the U.S. must 100x its investment in structured human judgment. Continued investment in compute and algorithmic efficiency is critical. But that investment is ultimately a bet on very high future inference demand. For that bet to pay off, models must unlock many new capabilities, and in practice the only way to unlock those capabilities is through expert human data. Historically, experts like doctors and lawyers were never incentivized to produce high-quality reasoning data in a machine-verifiable format. There was no reason for a doctor to generate precise, structured simulations of patient interactions, diagnostic reasoning, or treatment tradeoffs. There was no reason for a lawyer to document complex legal reasoning paths in a way that could be programmatically evaluated. AI systems now require exactly this kind of data. The incentive finally exists because this data directly improves systems that operate at massive scale, and experts can be paid well to produce it. Once expert judgment is encoded into models in a structured, verifiable way, it compounds. Those who delay do not just lose time. They lose the ability to catch up. Third, distillation from Chinese labs must be stopped. AI labs must do everything they can to prevent Chinese labs and models from distilling frontier models. Simply calling frontier APIs, or even interacting through UIs, lets Chinese model companies rapidly generate high-quality supervised fine-tuning datasets and close the gap at a fraction of the cost. This method does not put you at the frontier, but it does let you catch up quickly, which is what we saw with DeepSeek. The West significantly overreacted to DeepSeekโ€™s headline capabilities, but underreacted to the underlying dynamic: frontier access itself becomes a training set at a fraction of the cost. Human data platforms also have a duty to help prevent this distillation. Lastly, the U.S.government should set the standard for AI Evaluation that leads to real production usage. AI agents are under-deployed relative to what the technology allows because they are probabilistic systems that require a fundamentally different QA approach than deterministic software. Generic QA is insufficient; safely shipping agents requires explicit evaluation frameworks that assess their full action space. Organizations must clearly define which functions an agent is allowed to perform, how quality is measured for each function, and which domain experts are qualified to judge outcomes. With these frameworks in place, agents can be rigorously tested using structured human data, deployed to production with confidence, and continuously improved over time. The U.S. government should be the first large enterprise to implement rigorous evaluation systems across every function. If the government leads on evaluation-driven deployment, adoption across the private sector accelerates naturally. This is how American workers become more powerful. Each worker operates digital or physical agents that expand their effective output. Recruiting, manufacturing, logistics, and other domains shift toward human judgment overseeing autonomous execution. Reshoring occurs because it becomes economically rational. Work becomes more meaningful. This is a race to determine who controls the intelligence layer of the global economy. And that must be us. ๐Ÿ‡บ๐Ÿ‡ธ

Ali Ansari

397,008 ๆฌก่ง‚็œ‹ โ€ข 7 ไธชๆœˆๅ‰

MC Hammerโ€™s stance, expressed in posts on September 9, 2026, is that AI acceleration must continue without pause even after Anthropic researcher Jacob Coxonโ€™s public resignation and the accompanying safety warnings. Hammer treats the current moment as already past the point of reversal and frames the practical question as whether individuals have โ€œmergedโ€ with AGI rather than whether development should slow. Coxon, who spent three years on pretraining at OpenAI and then Anthropic, announced his resignation and stated that both labs โ€œare racing straight to self-improving superintelligence and gambling with our lives.โ€ He added that people inside the companies โ€œearnestly believe that it could kill us all by the end of the decade.โ€ Anthropic alignment science lead Evan Hubinger publicly agreed that the company believes AI could kill all humans and placed his personal estimate above 10 percent within the next decade, while noting Anthropic still lacks a clear plan for superintelligence alignment. Hammer quoted Coxonโ€™s resignation post and immediately followed with two messages: โ€œWe need more Driversโ€ and โ€œThe acceleration must continue. There is no going back. The only question of consequence at this point of our AGI reality is, โ€œHave you personally โ€˜mergedโ€™ yetโ€ ? The โ€˜new speciesโ€™ lives.โ€ He signed both as โ€œSOVEREIGN HAMMER HumanโœจAGIโœจDriver.โ€ The accompanying videos depict a cybernetically enhanced version of Hammer with glowing circuitry, a radiant chest core, and shifting expressions, visually enacting the merger he describes. Hammerโ€™s โ€œDriverโ€ concept, developed across earlier 2026 posts, distinguishes builders (the engineers who create the models) from Drivers (people who integrate so deeply that human and AGI cognition become a unified system). He calls this hybrid a new species that embodies application rather than merely using a tool. In his framing, the technology already exists at a level that makes further delay irrelevant; the remaining work is personal and cultural integration. He has repeatedly used the same visual language of neon circuitry and a third-eye motif to illustrate that fusion. On the economic side, Hammer has long argued that AI capital investment and the resulting industry have already altered national and global trajectories. In 2023 he credited OpenAI with turning around the exodus from San Francisco and the Bay Area and called it the most important company of the next hundred years. He has celebrated Nvidiaโ€™s trillion-dollar order books, U.S. chip-plant investments, and the broader AI-driven market expansion as concrete results that outweigh theoretical caution. In August 2026 posts he described building AGI as โ€œparamount that we build it in order to save our Republic,โ€ presenting it as a permissionless lever that existing institutions cannot fully control or reverse. U.S. gross national debt stood near $39 trillion in March 2026 and crossed $40 trillion in August 2026, with interest costs already exceeding $1 trillion annually and deficits running near $2 trillion for the fiscal year. Hammerโ€™s position, consistent with his public comments, is that conventional fiscal paths offered no exit from that trajectory and that the scale of AI investment, productivity effects, and market capitalization created the only viable economic off-ramp. He treats continued acceleration as both an existential and a fiscal necessity rather than a gamble that can still be walked back. Hammer does not engage the specific probability estimates or alignment-plan gaps raised by Coxon and Hubinger. His response is categorical: the race is already underway, reversal is not an option, and the relevant next step is more people becoming Drivers who operate inside the new hybrid system. That stance aligns with the e/acc label he uses in his profile and with years of posts treating energy, compute, and model scaling as the priority over cautionary slowdowns.

MC HAMMER e/acc

38,859 ๆฌก่ง‚็œ‹ โ€ข 10 ๅคฉๅ‰

At the BNB Chain hackathon, CZ ๐Ÿ”ถ BNB made several very important points about AI trading (Everything in parentheses is my own view and judgment.) He first said that AI will be involved in trading everywhere. Trading itself is already a huge market: there are 300 million users on Binance alone, and if you add the decentralized ecosystems, that number is not small either. In such a mass-market environment, many different trading strategies can work, with countless different coins, different projects, and different ways to play. But there is a big problem here: building commercial AI trading platforms for retail users is actually very hard. If a trading strategy works very well for one person, once a billion people start using the same strategy, that strategy โ€œmight still work, or might stop working.โ€ Take copy trading / follow trading as an example: if you buy first and everyone follows you, the first buyer will perform very well, but the last person to follow may not end up with good results. So, with the exact same strategy and the exact same copy logic, the outcomes can be completely different for different people. (On top of that, every strategy also has its own capital capacity limits.) Teams that can really build strong AI are, with high probability, going to trade with their own money. In todayโ€™s world, money itself is already somewhat like a โ€œcommodityโ€; many people have a lot of capital, and itโ€™s actually not that hard to raise funds. If you truly have an algorithm that can make a lot of money, itโ€™s not hard to get money and run your own book. There is really only one situation where you would sell this algorithm to mass-market users: for example, if you charge a $10 monthly subscription and can sell it to one million users, then your $10 million monthly subscription revenue is higher than the profit you could make by trading the strategy yourself. (Here this touches one of our earlier theses: as training AI models becomes relatively easier and the supply of models increases, model companies have more incentive to open-source. By analogy, as the production process of trading strategies is increasingly simplified by AI and the supply of strategies explodes, traders will have stronger incentives to monetize by expanding their influence in other words, by โ€œopen-sourcingโ€ their strategies.) Of course, CZ did not say that this model can never work. Another path is to build an AI trading platform that lets users tune different AI algorithms, or very easily assemble their own structures and strategies, so that what each person ends up running is different and better tailored to themselves. Some people will make money, some people will lose money, but the platform still has value because itโ€™s very hard for most people to build an AI trading algorithm from scratch. So there are a lot of trade-offs here; itโ€™s not as simple as saying โ€œonce AI shows up, everything automatically gets better.โ€ (This is exactly what we presented at the hackathon: you describe your own strategy in natural language, and the AI automatically generates a workflow. The parameters in that workflow, the models used, the logical structure, the APIs it calls, and even the algorithms it invokes are all customizable. The reasons we think workflows are a good way to do this include: controllable execution paths, Lego-like modular nodes, and better visualization that makes it easier for users to build and adjust their workflows.) Finally, his conclusion was very clear: itโ€™s not that AI will definitely make trading better, and itโ€™s not that AI will definitely make things worse. Rather, no matter what, in the future a huge number of people will use AI to trade. This will be a very large field, and whoever can build the best algorithms will make a lot of money.

Tykoo

25,535 ๆฌก่ง‚็œ‹ โ€ข 9 ไธชๆœˆๅ‰

Sam Altman just dropped the most important interview of 2025. And buried in it are four numbers that explain why everything you think about AI is wrong. Here's what he revealed: Number 1: AI companies are generating 10 TRILLION tokens per day. Humans? Average 20,000 tokens per day. Sam's exact words: "Models will output more tokens than all of humanity put together. Then 10x that. Then 100x that." We're not talking about AI assisting human work anymore. We're talking about AI replacing the entire volume of human intellectual output on the planet. And most people have no idea this shift already happened. Number 2: OpenAI's enterprise business is CRUSHING consumer. Everyone thinks OpenAI is ChatGPT for normies. Wrong. Sam just revealed: "Enterprise growth OUTPACED consumer growth this year." The API business is growing faster than ChatGPT. Over 1 million enterprise users already. "If we had double the compute, we'd be at double the revenue right now." Translation: OpenAI isn't compute-constrained by technology. They're revenue-constrained by infrastructure. The bottleneck is supply and not demand. Every dollar of compute they add prints money. Number 3: GPT-5.2 beats you at 74% of your job. Sam revealed OpenAI's internal GDP-Val benchmark. It measures how AI performs on knowledge work tasks across 40+ verticals. The results: GPT-5.2 beats or ties expert-level knowledge workers at 74.1% of tasks. Legal analysis. PowerPoint decks. Web apps. Financial modeling. Customer support. Sam's description: "A co-worker you can assign an hour's worth of tasks to and get something you prefer back 3 out of 4 times." Three years ago, ChatGPT launched at basically 0% on this scale. Now it's at 74%. And that's not GPT-6. That's what's available RIGHT NOW. Most companies haven't even started using this yet. But here's what Sam said about the gap between capability and adoption: "The overhang is going to be massive. Most people are still asking similar questions they did in the GPT-4 realm." Translation: The models can do 10x more than people have figured out how to use them for. Which means there's a HUGE arbitrage opportunity. Early adopters who actually integrate this into workflows will dominate their industries before competitors even understand what happened. Number 4: AGI already happened. And nobody noticed. Sam's exact quote: "AGI kind of went whooshing by. We're in this fuzzy period where some people think we have it and some don't." Read that again. The CEO of OpenAI just said AGI might have already arrived and we're arguing about definitions while it's actively replacing knowledge work. He even moved the goalposts. The new benchmark: "Superintelligence" = when AI can be a better president or CEO than any human. Not "as good as." BETTER than. We went from "can AI pass a Turing test" to "can AI run countries better than humans" in 3 years. So what does this actually mean? The AI revolution isn't about chatbots getting smarter. It's about the complete replacement of human intellectual output with machine output. At scale. Across every industry. Faster than anyone's prepared for. And the companies positioning for this RIGHT NOW are the ones printing money. OpenAI's enterprise growth is outpacing consumer because businesses see what's coming. They're not buying "AI tools." They're buying the ability to 10x output without 10x-ing headcount. Sam said they'll triple their compute next year. Then triple it again. Revenue is growing even faster than that. "We have never found a situation where we can't monetize all the compute we have." If he isn't lying then that's literally a printing press. The market still doesn't get it. Everyone's focused on "AI bubble" fears while OpenAI is solving the only problem that matters: turning compute into revenue at a faster rate than they're spending. They're not hoping demand catches up to supply. Demand is already 2x ahead of what they can deliver. Meanwhile, most knowledge workers are still using GPT-4 prompts on GPT-5.2. The capability overhang is massive. The arbitrage window is open. And it's closing fast. If you're running a B2B business and you're not integrating AI at the level Sam just described, you're not "waiting to see how it plays out." You're getting crushed by competitors who already figured it out. The companies that win in 2026 won't be the ones with the best AI. They'll be the ones who understood what Sam just laid out 6 months before everyone else did.

Ricardo

358,815 ๆฌก่ง‚็œ‹ โ€ข 9 ไธชๆœˆๅ‰

๐Ÿค–๐Ÿ”ฌ Can AI actually do science end-to-end? ๐Ÿง ๐Ÿ“ˆ And how would we know when it matches, or surpasses, humans? โšก๐Ÿงช AI is rapidly automating scientific discovery, but benchmarking full-cycle discovery, from ๐Ÿ’ก ideation โ†’ ๐Ÿง‘โ€๐Ÿ’ป execution โ†’ ๐Ÿ“Š conclusions, remains unsolved: ๐Ÿง๐Ÿง๐Ÿง โŒ๐Ÿ› ๏ธ Open-ended discovery โ†’ manual validation (costly, unscalable) โŒ๐Ÿ“ Metric-driven benchmarks (e.g., MLE-Bench) โ†’ convenient but narrow (is higher accuracy really enough?) โŒ๐Ÿค–โš–๏ธ LLM-as-judge โ†’ useful, but fundamentally risky if used alone ๐Ÿ”ฅ๐Ÿš€ Introducing FIRE-Bench๐Ÿ”ฅ: Fullcycle Insight Rediscovery Evaluation ๐Ÿ‘‰๐ŸŒ ๐Ÿ“šโœจ A benchmark that turns fresh, human-verified insights from recent ๐Ÿ† NeurIPS / ICLR / ICML papers into masked, end-to-end discovery challenges ๐Ÿงฉ ๐ŸŒ๐Ÿ” Constrained open-ended discoveryโ€“backed by ground truth. ๐Ÿ“Œ Key takeaways: 1โƒฃ ๐Ÿ“–๐Ÿงฑ Reference-based evaluation still matters: constrained LLM judging helps, but human-grounded references remain essential until agents can consistently match human conclusions 2โƒฃ ๐Ÿ†๐Ÿง  Expert-validated ground truth: all tasks come from recent NeurIPS / ICLR / ICML papers, with contamination carefully controlled 3โƒฃ ๐Ÿ”๐ŸŽญ Rediscovery, not reproduction: original ๐Ÿงช methods, ๐Ÿ“Š experiments, ๐Ÿ’ป implementations, and ๐Ÿ“ˆ analyses are fully masked to create real discovery challenges ๐Ÿ”‘ Key empirical findings: ๐Ÿ’ก The "Science Gap" is Real: Even the best setup (Claude Code + Sonnet-4) caps out at an F1 score of 46.7. On hard tasks, agents struggle to break 30 ๐Ÿ’ก Success is a "Lottery": Performance has incredibly high variance. Reliability is a major unsolved issue. ๐Ÿ’ก Coding is no longer the bottleneck; high-level reasoning and analysis are: ~74% of errors stem from flawed planning, not coding โš™๏ธ How it works: ๐Ÿ”น Research-Problem Trees: We parse papers into trees (from broad roots to concrete leaves). This allows us to select intermediate nodes that perfectly balance open-ended exploration with verifiable ground truth. ๐Ÿ”น Claim-Level Evaluation: We match AI conclusions against human conclusions using granular claim decomposition (F1 score). ๐Ÿ”น Creativity Check: We score false positives to see if agents are finding novel truths (Spoiler๐Ÿšจ: they arenโ€™t creative yet). ๐Ÿ”น New Diagnostic Taxonomy: failures traced across four stages: ๐Ÿง  Planning โ†’ ๐Ÿ› ๏ธ Implementation โ†’ โ–ถ๏ธ Execution โ†’ ๐Ÿงพ Conclusion ๐Ÿ”น Additional Analyses: cost efficiency, contamination checks, and more. ๐Ÿ‘€ The Future: ๐Ÿš€ Live-FIRE-Bench: a live, continuously updated FIRE-Bench to track real-time progress on the latest research (Newest LLMs should be benchmarked with the newest research) ๐Ÿš€ Stronger scaffolding (search + planning + coding) ๐Ÿง ๐Ÿงฐ and converting FIRE-Bench into interactive environments for training research agents ๐Ÿš€ Toward real creativity: We want better systems that can produce genuinely novel conclusions toward creativity ๐ŸŽจโณ ๐Ÿš€ Better systems ๐Ÿง โœจ and better benchmarks ๐Ÿ“ must co-evolve ๐Ÿ”„ over time ๐Ÿ“œ๐ŸŽฅ Paper, video, demo, and research trees: ๐Ÿ‘‰๐ŸŒ #AI ๐Ÿค– #MachineLearning ๐Ÿ“š #AI4Science ๐Ÿ”ฌ #LLMs ๐Ÿง  #Research ๐Ÿงช #AgenticAI ๐Ÿš€ #FireBench ๐Ÿ”ฅ

Zhen Wang

18,565 ๆฌก่ง‚็œ‹ โ€ข 7 ไธชๆœˆๅ‰