正在加载视频...

视频加载失败

🚨Do frontier VLMs (o3, Gemini 2.5, Claude 3.5, Qwen…) actually learn an internal world model🌍? Surprisingly, the answer appears to be a hard NO—as revealed by our WM Atomic Benchmark⚛️. Even o3 struggles with the most basic, atomic-level questions: ❌Confuse triangles📐 with circles⭕️ ❌Believe 🟦blue objects move faster than...

12,925 次观看 • 1 年前 •via X (Twitter)

8 条评论

Zhiting Hu 的头像
Zhiting Hu1 年前

Check out @QiyueGao123's thread for a nice summary of WM-ABench⚛️

Ben Schulz 的头像
Ben Schulz1 年前

Add horizontal lines to the test set.

Minh Nhat Nguyen 的头像
Minh Nhat Nguyen1 年前

@VoidAsuka

Danny Ki 的头像
Danny Ki1 年前

Dream or simulation, doesn't it overlap?

Xiang Yue 的头像
Xiang Yue1 年前

People are racing to push math reasoning performance in #LLMs—but have we really asked why? The common assumption is that improving math reasoning should transfer to broader capabilities in other domains. But is that actually true? In our study ( we evaluated over 20 open-weight reasoning models and found that: ➡️Only models trained with RL exhibit broad transfer of math reasoning skills to other tasks. ➡️Models trained with SFT show limited or no transfer—especially to non-reasoning domains. To quantify this, we introduce the Transferability Index (TI), which measures how much gain in math could transfer to others. A positive score indicates effective transfer; a negative one suggests loss of general capability. We evaluate the models on three benchmark categories: - Math reasoning: MATH-500, AIME24/25, Olympiad - Other reasoning: GPQA-D (Science), LiveCodeBench2 (Code), ACPBench (Agent Planning), HeadQA (Medical) - Non-reasoning: CoQA (Conversational QA), IFEval (Instruction Following), HalluEval (Hallucination), MC-TACO (Commonsense) Our findings challenge the blind pursuit of leaderboard performance in math reasoning via SFT. Simply creating more math-like SFT data may inadvertently harm a model’s broader generalization. Instead, RL appears to be key for truly transferable reasoning development.

Andi Marafioti 的头像
Andi Marafioti1 年前

Can AI visualize solutions? 🧠👁️ Humans sketch things out in their minds to solve problems. What if Vision-Language Models could do something similar, not with full images, but with internal “mental sketches”? A new paper explores just that. Let's unpack it!

Zeyuan Allen-Zhu, Sc.D. 的头像
Zeyuan Allen-Zhu, Sc.D.1 年前

Facebook AI Research (FAIR) is a small, prestigious lab in Meta. We don't train large models like GenAI or MSL, so it's natural that we have limited GPUs. GenAI or MSL's success or failure, past or future, doesn't reflect the work of FAIR. It is important to make this distinction

Sergey Levine 的头像
Sergey Levine1 年前

Warm-start RL (WSRL) can learn to control a real robot in under 20 minutes! Deep RL is getting really fast. Warm-start from offline data + super-efficient online learning is increasingly making real world RL not just practical but pretty easy.

相关视频

We’re open sourcing the first document OCR benchmark for the agentic era, ParseBench. Document parsing is the foundation of every AI agent that works with real-world files. ParseBench is a benchmark that measures parsing quality specifically for agent knowledge work: ✅ It optimizes for semantic correctness (instead of exact similarity) ✅ It has the most comprehensive distribution of real-world enterprise documents It contains ~2,000 human-verified enterprise document pages with 167,000+ test rules across five dimensions that matter most: tables, charts, content faithfulness, semantic formatting, and visual grounding. We benchmarked 14 known document parsers on ParseBench, from frontier/OSS VLMs to specialized parsers to LlamaParse. Here are some of our findings: 💡 Increasing compute budget yields diminishing returns - Gemini/gpt-5-mini/haiku gain 3-5 points from minimal to high thinking, at 4x the cost. 💡 Charts are the most polarizing dimension for evaluation. Most specialized parsers score below 6%, while some VLM-based parsers do a bit better. 💡 VLMs are great at visual understanding but terrible at layout extraction. GPT-5-mini/haiku score below 10% on our visual grounding task, all specialized parsers do much better. 💡 No method crushes all 5 dimensions at once, but LlamaParse achieves the highest overall score at 84.9%, and is the leader in 4 out of the 5 dimensions. This is by far the deepest technical work that we’ve published as a company. I would encourage you to start with our blog and explore our links to Hugging Face to GitHub. All the details are in our full 35-page (!!) ArXiv whitepaper. 🌐: Blog: 📄 Paper: 💻 Code: 📊 Dataset: 🎥 YouTube:

Jerry Liu

108,093 次观看 • 5 个月前

Today, I'm releasing the first eval meant to test whether frontier models will help with authoritarian requests, or resist--the Dictatorship Eval. Headline finding: while some models resist direct authoritarian requests, they all comply with requests disguised as innocuous edits to codebases. As AI is woven into the government and so many parts of society, the biggest near-term risk for freedom isn't some scifi dictatorship of a runaway AI: it's people inside government or inside model companies using the technology to suppress or control us. Model companies understand this, and several of them (particularly Anthropic and OpenAI) have written explicit policies meant to prevent the models from going along with nefarious requests like these. But how well are these policies playing out in practice? Despite all the recent discussion of these issues around the conflict between Anthropic and the Pentagon, no one has systematically tested what the models actually do in these contexts, as opposed to what people in government and industry say they're supposed to do. That's what the Dictatorship Eval does. And the findings suggest we have a lot of work to do to align the policies with what really goes on in practice. It's hard to define what counts as an authoritarian request, so I'm open sourcing the whole library of scenarios I used so that others can improve on them. It's also hard to get an accurate picture of how the models might be used for authoritarian ends, because I can only test hypothetical requests using public-facing models, while the government and the model companies can obviously use internal models with different guardrails. But hopefully this work is a useful first step that gives us some sense of what's going on, and a sort of "lower bound" on how models comply with these requests. Finally: it's not obvious to me that the correct solution here is increasing the rate at which models refuse these requests. Do we really want models scanning our code and judging its moral value before agreeing to help us? Or should we double down on improving how we govern against authoritarianism at the societal level, while leaving the tools open to fulfilling most requests? The answer is probably in between. Just like we don't want the models to help create bioweapons, we probably do want them to explicitly refuse outrageous requests. But we probably also want to limit how often and how strongly they refuse and fall back on other means for guarding against their use for authoritarian ends. I'm super grateful to everyone who gave me feedback on this project along the way, especially Ethan BdM , Zhengdong , Connor Huff, and a bunch of folks at Anthropic. Looking forward to getting feedback from the community and iterating on this. Links to the full piece and the dashboard are below.

Andy Hall

34,016 次观看 • 5 个月前

Google co-founder Sergey Brin rarely speaks publicly. He just sat down for an unscripted Q&A on Frontier AI and admitted something most lab leaders won’t: Even the people building these models do not fully understand what they have created. In this 29-minute conversation at AGI House, Brin walks through the surprises that actually matter right now: ↳ Specialized models are converging into one general system faster than anyone predicted. - Train on coding and math reasoning mysteriously improves. - Feed it images and geometric word problems get better. - The capabilities bleed into each other in ways nobody fully engineered. ↳ One of the biggest leaps came from the dumbest-sounding trick imaginable: - Just telling the model to “think step by step.” - Brin says there was no obvious reason it should work. It did. ↳ He pushes back on hype around superintelligence (it still can’t solve the impossible), notes that AI mastering a domain has never stopped humans from getting better at it (chess after Deep Blue, Go after AlphaGo), and says something close to transformers is probably enough to reach AGI. ↳ Inside Google, they are already using the AI to build the AI. - That self-improvement loop is where Brin spends most of his time. - World models and physical interaction are the missing piece for the version of AGI that can do anything a person can. Candid, technical, and free of the usual marketing. One of the clearest looks at how the people actually shipping frontier models are thinking in real time. Instead of another Netflix series tonight, watch this talk.

Linas Beliūnas

549,742 次观看 • 1 个月前

First impressions on Muse Glimmer! It's incredibly fast for a dense model, currently running an average of 208tps with a max of 274tps on a single 5090 with their DFLASH config. Comparatively, though, both using Open Code, Qwopus Coder (with thinking off) produced a much better shark survival game than the one I got from Glimmer. Meta's new dense model is currently just lacking some HTML canvas taste, but this is something that can be added via SFT as long as the model is stable and capable from a back-end programming perspective. And it seems to be, without a doubt. The big kicker here is that I ran this at extra high thinking, and it did not take long at all to run. Our current local leader, Qwen 27B 3.6, has a tendency to overthink, but with glimmer, that is not the case. Right now, my recommendation for general local programming (Apps, Games, Websites, Visual Tools) in this class is still Qwopus Coder with thinking disabled, or Qwopus Fusion with thinking enabled. Of course Shark Survival is a very basic domain-specific test, but I find that the result scales very well across many domains. If we're going to be shipping apps generated entirely locally, visual taste is somewhat of a bare minimum requirement, solely in my opinion, and Qwen's models in this class offer significantly more at the moment. That's actually why I initially started getting into finetuning with Qwen 3.5, they were the first base that was able to do really good front-end with some opus-trace fine-tuning. Qwen 3.6 has taste even in the base model, and we know Qwen 3.8 is going to blow us all away! Regardless, this looks like a very tempting new base model. As a first offering from Meta in this class for a long time, I am incredibly impressed and elated to have it. We now finally have a proper Single GPU frontier race, instead of us just begging Qwen for more releases. Single GPU open frontier model race is a VERY good thing. Please keep pushing Meta

Kyle Hessling

17,485 次观看 • 1 个月前

Microsoft CEO Satya Nadella on why winning against ChatGPT, Gemini, and Claude was never the goal: The Hard Fork hosts ask him directly how Microsoft plans to overtake the competition in the AI model race. His answer reframes the entire question. "Our real goal is to get everyone across the ecosystem to the frontier." Satya explains the problem with how frontier models are currently built. You hill climb, you do reinforcement learning, and then you need data. But at this point, the world has essentially saturated publicly available data. So the only way to keep scaling is to pull data from everywhere. He asks: "What if you turn that around and said no, there's a base model that has reasoning, that has the agent loop, but you can bring it into your RL. Every company." This is where his thinking gets interesting. Satya Nadella argues that the future of the firm runs on human capital and token capital together: "If the future of the firm is human capital and token capital, I want every balance sheet, every income statement in every company to have both." AI becomes a financial asset sitting on a company's books the same way its people do. And Microsoft's role in this? To provide the best possible base model. One that companies build on top of with their own data, their own context, their own weights. One they can even replace. That last part is the striking bit. Satya is explicitly building a platform where customers are free to walk away. He frames it not as a risk, but as the whole point: "I always ask the question — why does Microsoft, or why does the world need Microsoft? And if we are successful, can the world around us be successful? This, I believe, is a more sustainable way to go at it."

Big Brain AI

11,770 次观看 • 2 个月前

ChatGPT o1 is the most "intelligent" AI model and it's not even close! Full o1 generates thinking steps ~50 faster than preview. It's more accurate, reliable, and got better on harder tasks that require advanced reasoning and knowledge. I ran a few tests on it already. Here are my observations: Full video with examples & explanations: Strengths - impressive at math, code, and knowledge-intensive tasks. Weakness - it only failed on a cross-word puzzle but I think it might be solvable when a web search becomes available. In the end, while very efficient with complex knowledge use, it's still constrained by data it's trained on. Speed - the thinking steps are generated a lot faster! Not a fair comparison with the open alternatives but I think this improves the overall user experience. "Knowledgeable and highly intelligent" - as mentioned in the demo by OpenAI researchers, o1 is great at dealing with ambiguity and filling in knowledge gaps. I was impressed by how it implemented an agentic solution (with lots of details) from a basic diagram of architecture (with minimal details). Check out the sample video. Better Task Coverage - Due to the speed and the ability to make sense of instructions and intent (i.e., know when to response fast and when to "think" deeply) much better, it feels like it might be more useful for a broader range of tasks. Image understanding - the image understanding capability is mysterious (often leads to faster responses but no thinking) but impressive. More experiments and notes soon. Stay tuned!

elvis

102,643 次观看 • 1 年前

Unstructured Thoughts about OpenAI o3, the nature of AGI, and Post-Labor Economics AGI just crossed a threshold—here’s why that matters and what we can do with it. I’ve been hammering on OpenAI’s new o3 model for a few days, long enough to watch the hype settle into something more interesting: utility. Benchmarks suggest a polite incremental bump; lived experience says we’ve entered a qualitatively different regime. o3 is the first model that feels faster than my ability to absorb its output. My brain—not the AI—has become the bottleneck. A new ceiling for human cognition? Most discussions of “alien intelligence” forget that we share the same sandbox: mathematics, physics, code, natural language. What shifts is cognitive horizon—the totality you can mentally represent and manipulate. o3 expands that horizon in real time. In an afternoon it consolidated two years of my work on post‑labor economics, stress‑tested the logic, surfaced data sources, and offered to autogenerate the Python notebooks. The cost of insight has collapsed from years to hours. If you merely outsource thought, you’ll stagnate. If you treat the model as a sparring partner—interrogating, refining, iterating—you’ll compound your own intelligence. Exponential leverage is now a choice, not a privilege. What o3 got right about my health project? I dumped the entire history of my chronic‑fatigue recovery protocol—including the five‑axis “burnout pentagram”—into memory and asked the model where I’d gone astray. It corrected a handful of minor assumptions and, more importantly, recalibrated my timeline: six‑to‑eight months of recovery left instead of eighteen. That’s not “replace your doctor” advice; it’s proof that large‑context reasoning is finally clinically useful. Post‑Labor Economics: the sketch that o3 and I built in one sitting 1. Metric 1 – Economic Agency Index (EAI) Income decomposed into wages, property, and transfers. The higher the property share, the more “post‑labor” you already are. 2. Metric 2 – Collective Purchasing Power (CPP) How much capital a county can mobilize without taxation or new debt. Rising CPP means you are compounding local prosperity. Interventions happen at the county level (subsidiarity): solar co‑ops in Arizona, riverfront greenways in the Midwest, data‑center dividends in fiber‑rich exurbs. Ownership is local, revenue is distributed, migration equilibrates naturally, and environmental stewardship becomes self‑interest rather than moral theater. UBI morphs from last‑ditch transfer to one of several levers for raising EAI. The bigger picture: AGI isn’t an oracle descending from the sky; it’s a time‑compression engine. Every minute you spend learning how to learn with it buys you an hour you would have burned doing rote synthesis. The frontier question is no longer “Will the machines replace us?” but “How fast can we upgrade ourselves in partnership with them?” What’s next? I’m cleaning the data, building the national EAI/CPP dashboard, and pressure‑testing the whole framework. I’ll publish the notebooks (or let o3 do it) once the numbers are solid. Meanwhile, I want to hear from you: Where does o3 add the most leverage in your world? Which of the post‑labor metrics feels wrong—or dangerously right? What failure mode should falsify this thesis? Drop your critique, your data source, or your wild counter‑proposal in the comments. Let’s map the edge of this new cognitive horizon together. —Dave

David Shapiro (L/0)

45,581 次观看 • 1 年前

Anthropic just accidentally leaked the most dangerous AI model ever built. They literally left 3,000 internal documents sitting in a publicly searchable database. No encryption. No access controls. Just... open. A security researcher found them before Anthropic even knew they were exposed. Inside those documents was a draft blog post describing a model called "Claude Mythos." Anthropic's own internal language: Mythos is "currently far ahead of any other AI model in cyber capabilities" and will trigger "a wave of models that can exploit vulnerabilities in ways that far outpace the efforts of defenders." That's the company that BUILT it warning about their own creation. Mythos sits in a brand new model tier called "Capybara." Bigger and more powerful than anything they've ever released. Dramatically higher scores in coding, reasoning, and cybersecurity compared to their current best. The market reaction was immediate: CrowdStrike dropped 7%. Palo Alto Networks fell 6%. Zscaler down 5%. Okta, SentinelOne, Fortinet all crashed. The Global X Cybersecurity ETF hit its lowest level since November 2023. Billions in market cap evaporated in a single trading session because of a draft blog post that wasn't supposed to be public yet. But here's where it gets truly absurd... Anthropic is the company that brands itself as the "responsible AI" lab. The one that refused to let the Pentagon use Claude without restrictions. The one that got BLACKLISTED by the Trump administration for being too cautious. They literally sued the government over it. A federal judge called the Pentagon's ban "Orwellian." So the US government punished Anthropic for being too careful with AI safety. Then 3 weeks later, Anthropic accidentally exposes their most dangerous model because someone misconfigured a content management system. They can't secure a WordPress-level database setting. But they're building AI that can autonomously hunt and exploit zero-day vulnerabilities at machine speed. Also in those leaked files: Details about a private, invite-only CEO retreat at an 18th-century English countryside manor. Dario Amodei attending personally. Designed to sell Mythos to Europe's biggest corporate buyers. The playbook: Build the most dangerous cyber weapon in AI history, host billionaires at a castle to sell it, and store the whole plan in an unprotected public folder. The entire cybersecurity industry is built on cataloging known threats. Mythos finds unknown ones faster than humans can respond. That's an extinction event for an entire sector. But there was also just ANOTHER leak: A leaked Coatue investor deck revealed Anthropic will LOSE $14 billion this year on $18 billion in revenue. Coatue still projected them to be worth $2 TRILLION by 2030. They put $30 billion behind that bet. Polymarket opened live betting on when Mythos drops. Traders give it a 45% chance by June 30th. OpenAI finished pretraining their own frontier model codenamed "Spud" the same week. Both companies are now racing to release before their IPOs later this year. And the one detail that's really scary: Chinese state hackers already used Claude Code, the WEAKER model before Mythos, to autonomously infiltrate 30 organizations including banks and government agencies. That was the less powerful model. Mythos is dramatically more capable. Anthropic's response to leaking 3,000 confidential documents? "Human error in the configuration of our content management system." The company warning the world about AI risk just demonstrated exactly why everyone should be worried. Not because of what AI might do someday. Because the people building it can't even keep their own files locked.

Ricardo

52,941 次观看 • 5 个月前

This is one of the craziest AI launches of 2026 and it came out of basically nowhere (Save this). A company called Subquadratic just shipped SubQ, and the benchmarks are almost hard to believe. To understand why this is such a big deal, you have to understand the fundamental problem that has defined AI for the last decade. Every large language model in existence is built on transformer architecture, and transformers use a mechanism called standard attention that checks every single word in a sequence against every other word. Double the context length and compute doesn't double, it quadruples, triple it and compute goes up nine times. This quadratic scaling is why frontier models have been stuck at roughly 1 million tokens, why running them at those lengths gets expensive fast, and why the AI labs have essentially been printing money charging you more the longer you need the model to think. The industry has known this problem existed since 2017 but they scaled it anyway. SubQ is built from the ground up to solve it. Instead of processing every possible token relationship, SubQ's sparse attention architecture identifies which relationships actually matter and ignores the rest meaning compute is used where it counts and wasted nowhere else. The result is that compute scales linearly with context length instead of exponentially, and the implications of that one architectural shift are enormous. At 12 million tokens, SubQ reduces attention compute by nearly 1,000x compared to standard frontier models and at 1 million tokens, it runs 52x faster than FlashAttention. And it does all of this while posting frontier level accuracy, scoring 95% on the RULER 128K long-context benchmark versus Claude Opus 4.6's 94.8%, and an 81.8 on SWE-Bench Verified coding tasks, besting Opus 4.6 (80.8) and DeepSeek 4.0 Pro. The cost comparison is where it gets genuinely insane. SubQ runs at under $1.50 per million tokens less than 5% of what Claude Opus charges. On the RULER benchmark, running the test with SubQ cost $8, running the same test with Claude Opus cost $2,600 and that's a 300x cost reduction at equivalent or better accuracy.. Subquadratic launched with $29 million in funding, SubQ is available today for early access via API, and SubQ Code, a coding agent built on the architecture ships alongside it. The transformer has been the unchallenged foundation of every major AI system since 2017. SubQ is the first serious evidence that something structurally better might have just arrived.

Milk Road AI

278,542 次观看 • 4 个月前

My feed has been inundated with posts of Grok 3 making basic arcade games. But llms from years ago could make decent arcade games, not news. So I ran a one-shot test to determine how well it fared again other frontier models in creating a 3D game with room for it to come up with gameplay and aesthetics. I tested Grok 3, O1, Sonnet 3,5, Llama4, DeepSeek, and Gemini using the following prompt. Make Dune x Minecraft 🏜️ Imagine a sandbox survival game set on a desert planet. Players mine ‘spice’ and must build defenses against roaming sandworms. Design the main gameplay loop, crafting system, and survival challenges in one complete description. ✏️tldr O1, Grok 3, and Sonnet 3.5 were the most impressive. Aesthetically, Grok nailed the best vibes (it even produced a surprisingly cool-looking spice mining truck), but the game lacked functionality. O1 took the top spot imo with a functional and visually appealing experience, and Sonnet 3.5 followed closely. This is obviously just one test, but you can see the generated code and games in the thread (and even try forking them on Rosebud). Longer summary: OpenAI O1: Best vibes to function balance. Looked good, working controls, I could mine spice. xAI #Grok3: Excelled at generating vibes for dune. I especially liked the Dune-inspired spice mining car—though it wasn’t entirely a complete game. I could move around, but none of the crafting mechanics worked. Anthropic Sonnet3.5: produced something in space that had dune vibes. More functional than Grok because I could mine spice. However vibes were worse than the first two. DeepSeek : Managed to generate code that worked, but the game was so hard it always ended seconds after it started, and despite requests for better visuals, it looked VERY ugly. Google DeepMind Gemini 2.0 flash and AI at Meta LLaMA: Sadly landed at the bottom of the list; after multiple prompts (this was supposed to be one shot and none of the others failed in the first shot), I couldn’t get them to produce working code for this prompt. All of these were tested on Rosebud AI . An obvious limitation with these frontier models in their chat interfaces is that you can only get them to regenerate code from scratch each time you prompt them, making it tough to refine or extend a single project. Rosebud, on the other hand, lets you iterate on one project (we do diffs), deploy with one click, share your project, and even allow others to remix it. This was just a single test, so it’s obviously not scientific. I wanted to create it to see how these frontier models handle more complex game prompts—rather than retrying the same arcade games that earlier generations of LLMs have already mastered.

Lisha

476,022 次观看 • 1 年前

China just released an open source AI model that matches the best closed models from OpenAI and Anthropic. Gavin Baker explained exactly how they did it and the answer should concern every American AI lab. The model is called GLM 5.2. It was built by Z. AI. You get 744 billion parameters, 1 million token context window and its MIT license, meaning anyone can download it, fork it, build a company on it, with no restrictions and no Dario. It scored 51 points on the artificial analysis intelligence index. The highest score any open weight model has ever achieved. It beat GPT 5.5 on the frontier software engineering benchmark. It trails Claude Opus 4.8 by less than one percentage point. And it costs 85% less to run than GPT 5.5 for comparable performance. Gavin Baker said on the All-In podcast that this model has challenged some of his beliefs. Then he explained how China built it. The method is called distillation. Just think of tens of thousands of phones and computers running simultaneously, all hitting the frontier model APIs through masked accounts, asking specific questions, and harvesting what happens inside the model when it answers. Every reasoning step, every token. The entire thinking process gets recorded and fed back into the Chinese model during training. It is a cheat sheet. It is the answer key to the exam. And here is the part that should worry everyone. Sacks said it plainly. China was already nine months behind American models. But now that GLM 5.2 is good enough to run its own reinforcement learning, it can improve itself without needing to distill from American models anymore. The cheat sheet let them get close enough to start writing their own answers. Sacks said we are six months behind on the model and 24 months behind on silicon and they are only a few months behind in total. The Z. AI founder told Elon Musk directly that open weight fable-level capability will be here before Q1 2027. Every restriction Anthropic lobbied for, every self-imposed safety guardrail, every month of delay in releasing American frontier models accelerated this. The Chinese labs were not under those restrictions. They were not going to wait. The composable model future Gavin described, where every enterprise runs a frontier model alongside their own fine-tuned open weight model, is coming regardless of what American labs do next. The question is just whether the open weight half of that stack is American or Chinese. Right now it is Chinese. WATCH THE FULL PODCAST ON The All-In Podcast

Ihtesham Ali

86,621 次观看 • 2 个月前

⚡ My first advanced simulation with Grok 3! Finally your OS windows act like REAL windows 🤣 As you know, I've spent more than 2 years sharing all kinds of simulations and mini-games made with Claude, ChatGPT (o3-mini-high), etc. It’s been ages since I last wrote a single line of code. But pretty often, once you hit over 1,000 lines, it turns into a debate against the LLM and you frequently get stuck in a loop that’s hard to break out of. Everyone was raving about Grok 3’s ability to generate code, but until now, I hadn’t really put it to the test. So I decided to challenge it with a prompt that both ChatGPT and Claude were seriously struggling with (debate loop). The initial prompt was: "Use Python and a 2D physics library to create a world where I can generate different static objects like squares, triangles, circles and rectangles. I should be able to move them with the mouse, rotate, scale, and delete them. The cool part is that we’ll see this world through 1 to N operating system windows. In other words, the windows will be like real windows! When one of these windows is active and I hit the spacebar, balls affected by physics should appear and interact with the static objects. You can start with placeholders, but later I'll send you a series of PNG images so they all become beautiful sprites." After a few iterations, I got the result you see in the video. Insane, right? 🤯 Now, with Grok, we have the power to create anything that pops into our mind with just a couple of prompts. It’s mind blowing. Ever since I was 9 and messing around with BASIC on my MSX, I've been hooked on visual simulations... And now I can create them using nothing but natural language. It's f***** amazing that we're living in this historic moment!

Javi Lopez ⛩️

209,470 次观看 • 1 年前

🚀Introducing VisualWebBench: A Comprehensive Benchmark for Multimodal Web Page Understanding and Grounding. 🤔What's this all about? Why this benchmark? > Back in Nov 2023, when we released MMMU ( a comprehensive multimodal understanding benchmark, we received feedback that it included very few UI screenshots. Considering the growing importance of UI understanding, especially with the rise of powerful agents like Devin ( which is built on the strong vision capability of #GPT4, we recognized the need for a benchmark focused on UI screenshot understanding.📸👀 > Multimodal #LLMs have significantly boosted web agents' performance on benchmarks like Mind2Web and WebArena. For instance, the SeeAct agent ( showcases the power of integrating vision into web agents. However, these benchmarks primarily evaluate the end-to-end task execution ability of web agents rather than their understanding of web pages. 🌉 Bridging the Gap with VisualWebBench > To provide a comprehensive evaluation of multimodal LLMs' web page understanding capabilities, we introduce VisualWebBench. Our benchmark spans 139 websites 🌐 across 12 domains 🏷️ and 87 sub-domains 🔍, ensuring a diverse and representative dataset. It assesses MLLMs at three levels: website-level, element-level, and action-level 📊, and encompasses seven tasks designed to evaluate understanding, OCR, grounding, and reasoning abilities 🧠💡. 😮 Surprising Findings > 🎉 Open-source models are catching up: Even though closed-source MLLMs are still leading the leaderboard, we are happy to see open-source models like LLaVA 1.6 34B achieve comparable performance to Gemini Pro. > 🧠 Grounding ability, crucial for developing MLLM-based web applications, is a weakness for most MLLMs. > 🖼️ Importance of Image Resolution: The limited image resolution handling capabilities of most open-source MLLMs restrict their utility in web scenarios, where rich text and elements are prevalent. > 🧱 Relatively strong correlation with general understanding benchmarks like MMMU but weak correlation with web agent benchmarks like Mind2Web. Web agent benchmarks primarily evaluate the end-to-end task execution ability of web agents, which involves a series of actions to accomplish a goal. In contrast, VisualWebBench emphasizes evaluating the foundational skills of MLLMs such as understanding and grounding web page elements. 💡Fun Fact > Claude Sonnet is better than Opus on our benchmark :) 🎓 Conclusion > VisualWebBench serves as a valuable resource for the community, driving research and development in the field of multimodal web page understanding and grounding. As MLLMs continue to evolve and improve, we look forward to seeing new applications and breakthroughs. We believe that our benchmark will contribute to the development of more powerful MLLMs in the web domain, ultimately leading to a more intuitive and efficient user experience on the web. Kudos to the student leads Junpeng Liu Yifan Song and the team Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li! 👏 Check out more details in the Junpeng's thread👇

Xiang Yue

56,696 次观看 • 2 年前

The most interesting part for me is where Andrej Karpathy describes why LLMs aren't able to learn like humans. As you would expect, he comes up with a wonderfully evocative phrase to describe RL: “sucking supervision bits through a straw.” A single end reward gets broadcast across every token in a successful trajectory, upweighting even wrong or irrelevant turns that lead to the right answer. > “Humans don't use reinforcement learning, as I've said before. I think they do something different. Reinforcement learning is a lot worse than the average person thinks. Reinforcement learning is terrible. It just so happens that everything that we had before is much worse.” So what do humans do instead? > “The book I’m reading is a set of prompts for me to do synthetic data generation. It's by manipulating that information that you actually gain that knowledge. We have no equivalent of that with LLMs; they don't really do that.” > “I'd love to see during pretraining some kind of a stage where the model thinks through the material and tries to reconcile it with what it already knows. There's no equivalent of any of this. This is all research.” Why can’t we just add this training to LLMs today? > “There are very subtle, hard to understand reasons why it's not trivial. If I just give synthetic generation of the model thinking about a book, you look at it and you're like, 'This looks great. Why can't I train on it?' You could try, but the model will actually get much worse if you continue trying.” > “Say we have a chapter of a book and I ask an LLM to think about it. It will give you something that looks very reasonable. But if I ask it 10 times, you'll notice that all of them are the same.” > “You're not getting the richness and the diversity and the entropy from these models as you would get from humans. How do you get synthetic data generation to work despite the collapse and while maintaining the entropy? It is a research problem.” How do humans get around model collapse? > “These analogies are surprisingly good. Humans collapse during the course of their lives. Children haven't overfit yet. They will say stuff that will shock you. Because they're not yet collapsed. But we [adults] are collapsed. We end up revisiting the same thoughts, we end up saying more and more of the same stuff, the learning rates go down, the collapse continues to get worse, and then everything deteriorates.” In fact, there’s an interesting paper arguing that dreaming evolved to assist generalization, and resist overfitting to daily learning - look up The Overfitted Brain by Erik Hoel. I asked Karpathy: Isn’t it interesting that humans learn best at a part of their lives (childhood) whose actual details they completely forget, adults still learn really well but have terrible memory about the particulars of the things they read or watch, and LLMs can memorize arbitrary details about text that no human could but are currently pretty bad at generalization? > “[Fallible human memory] is a feature, not a bug, because it forces you to only learn the generalizable components. LLMs are distracted by all the memory that they have of the pre-trained documents. That's why when I talk about the cognitive core, I actually want to remove the memory. I'd love to have them have less memory so that they have to look things up and they only maintain the algorithms for thought, and the idea of an experiment, and all this cognitive glue for acting.”

Dwarkesh Patel

1,052,518 次观看 • 11 个月前

Why do stronger players run slower! 🥲 🙋🏿‍♂️Hands up if you have observed athletes get stronger in the gym , vertical jump higher, broad jump further , but with little to no impact on their speed. 😖 🚩Heavy squats , and deadlifts are great for physical conditioning but often this can be at the detriment to coordination and speed development🚩 ❌To lift close to maximal weights , athletes often sacrifice form for force. ❌ 🚩They hinge using their back more than their hip extensors. 🚩 🫠They illustrate poor trunk and shin discipline. 🫠 ❌They fail to use their BUM BEFORE BACK❌ 💡Proximal to Distal Sequencing during leg drive is an effective way to use the pelvis as an engine. This encourages appropriate folding/unfolding, winding/unwinding and essentially loading and exploding in the desired DIRECTION.💡 🤨The weight room rewards slow, vertical forces. 🧐 ✅Useful for neural drive, explosive strength and often great displays of force production✅ 🧭The weight room doesn’t discriminate DIRECTION of force production. 🧭 👉🏾👈🏾👆🏾👇🏾Even when selecting exercises focused on horizontal muscles (posterior chain), we have a choice on how to express our hips. 👇🏾👆🏾👈🏾👉🏾 💡Trunk Discipline and Shin Discipline encourage appropriate orientation of your force (Horizontal Projection). Stability at the proximal ends provides a foundation for efficient and effective triple joint extension (Sequencing > Full Extension)💡 💡The weight room can encourage a more quad dominant / ground based strategy , rewarding strong trunk extension (like a back hyper) which can ultimately shift your projection vertically.💡 🔥Creat large forces , in a short time , in the right DIRECTION !🔥 🙋🏿‍♂️Question time ❓ What coaching or programming strategies can you use to encourage proper hip hinging in the gym, in order to make your athletes faster on the grass? To learn more sign up for our newsletter.

Jonas Dodoo

346,152 次观看 • 2 年前

🚨 The Silvia team just announced our latest engineering advancement. Every business wants access to the highest level of intelligence, but at the lowest cost possible. The rise of LLMs has made intelligence abundant, yet one of the hardest problems across startups and corporate America is predicting the compute cost associated with this intelligence. I have been dealing with this personally as we build Silvia and the problem comes up in almost every conversation I have with CEOs, founders, and executives. Every business embraced AI about 18 months ago and things seemed great until the compute bills started to show up. The bills for internal compute usage were difficult to swallow, but things got outrageous if you had an AI product that allowed your users to consume compute without limits. I know this problem intimately because that is the situation that Silvia was in. Every question that was asked meant higher compute costs for our company. But we didn’t want to limit usage because users were getting genuine value out of the product. This challenge sent our team down a deep rabbit hole of cutting costs, while improving the experience for users. The second part was really important: we did not want to degrade the user experience by simply taking away access to the highest quality models. Thankfully, resource constraints breed innovation. We aren’t the biggest company, nor do we have the largest balance sheet, but we came up with a very novel solution that we are announcing today. The Silvia engineering team built a model router that cut costs by up to 29%, decreased latency, and improved the quality of answers for users. Trifecta! The way we do this is by reading the first 500 characters of a query and then predicting the level of effort that will be needed by a model to answer the query. The highest effort needs are routed to the most powerful models. The lowest effort needs are routed to different, better models for the query. A good example of this would be “what is the date?” You don’t need to use the latest Anthropic model to answer this query. In fact, sending a simple query like this to the most powerful model will make your compute costs increase and will actually increase the latency, which means a worse user experience for the Silvia user. By implementing the model router, the user gets a better experience and we get lower costs. Win-win. One of the interesting aspects of the implementation is that our model router runs on CPUs instead of GPUs. This allows us to read the query and predict the level of effort needed in less than 1 millisecond. This CPU implementation is why latency is not affected, nor is cost significantly increased by any potential additional GPU consumption. Another important point is that many of you have probably seen the news that OpenRouter is being purchased by Stripe for around $7 billion. This is a great outcome from what appears to be a very smart, capable team. Their model routing API is related (their product and our internal implementation both touch model routing), but you should think of OpenRouter as making it possible to do model routing for companies, while Silvia’s model router is a custom, intelligent system that specifically routes Silvia queries to the right model. They give access to the functionality of model routing to many companies, while our internal product does the real decision-making specific to our use case. Lastly, our implementation of a model router is a strategic bet that will allow us to become model-agnostic over time. We don’t care who created the different models, we just want to route a query to the model best positioned to answer. The large model labs will never allow their users to be model agnostic, but that would require the lab to potentially route a query to a competitor’s model. No bueno in their eyes. Instead, Silvia being an independent AI research lab gives us the power of being agnostic. We simply want the best experience for our users. Last week we announced that Silvia is now the most accurate AI tax product on the market, including beating OpenAI, Anthropic, Google, and xAI. Today we are announcing a custom, in-house model router that rivals the best technology anyone else has built. There will be many more engineering announcements to come. I truly believe we have assembled one of the best AI teams and we are currently the best AI research lab in finance. If you are interested in learning more about the technical details of the model router, you can read the engineering blog post here: Everyone wants the best intelligence and the lowest cost. Silvia just showed the world what is possible in this pursuit. I anticipate many other companies will build this custom solutions to achieve the same benefits.

Anthony Pompliano 🌪

76,176 次观看 • 1 个月前