正在加载视频...
视频加载失败
🚨Do frontier VLMs (o3, Gemini 2.5, Claude 3.5, Qwen…) actually learn an internal world model🌍? Surprisingly, the answer appears to be a hard NO—as revealed by our WM Atomic Benchmark⚛️. Even o3 struggles with the most basic, atomic-level questions: ❌Confuse triangles📐 with circles⭕️ ❌Believe 🟦blue objects move faster than... show more
12,925 次观看 • 1 年前 •via X (Twitter)
8 条评论

Check out @QiyueGao123's thread for a nice summary of WM-ABench⚛️

Add horizontal lines to the test set.

@VoidAsuka

Dream or simulation, doesn't it overlap?

People are racing to push math reasoning performance in #LLMs—but have we really asked why? The common assumption is that improving math reasoning should transfer to broader capabilities in other domains. But is that actually true? In our study ( we evaluated over 20 open-weight reasoning models and found that: ➡️Only models trained with RL exhibit broad transfer of math reasoning skills to other tasks. ➡️Models trained with SFT show limited or no transfer—especially to non-reasoning domains. To quantify this, we introduce the Transferability Index (TI), which measures how much gain in math could transfer to others. A positive score indicates effective transfer; a negative one suggests loss of general capability. We evaluate the models on three benchmark categories: - Math reasoning: MATH-500, AIME24/25, Olympiad - Other reasoning: GPQA-D (Science), LiveCodeBench2 (Code), ACPBench (Agent Planning), HeadQA (Medical) - Non-reasoning: CoQA (Conversational QA), IFEval (Instruction Following), HalluEval (Hallucination), MC-TACO (Commonsense) Our findings challenge the blind pursuit of leaderboard performance in math reasoning via SFT. Simply creating more math-like SFT data may inadvertently harm a model’s broader generalization. Instead, RL appears to be key for truly transferable reasoning development.

Can AI visualize solutions? 🧠👁️ Humans sketch things out in their minds to solve problems. What if Vision-Language Models could do something similar, not with full images, but with internal “mental sketches”? A new paper explores just that. Let's unpack it!

Facebook AI Research (FAIR) is a small, prestigious lab in Meta. We don't train large models like GenAI or MSL, so it's natural that we have limited GPUs. GenAI or MSL's success or failure, past or future, doesn't reflect the work of FAIR. It is important to make this distinction

Warm-start RL (WSRL) can learn to control a real robot in under 20 minutes! Deep RL is getting really fast. Warm-start from offline data + super-efficient online learning is increasingly making real world RL not just practical but pretty easy.
