
Wenhu Chen
@WenhuChen • 25,649 subscribers
MSL@Meta. I led PoT, MMMU, MMLU-Pro, MAmmoTH, General-Reasoner, VL-Rethinker, Pixel-Reasoner. I contributed to Gemini-2.5. Prev @GoogleDeepMind.
Videos

🚀 New Paper: Pixel Reasoner 🧠🖼️ How can Vision-Language Models (VLMs) perform chain-of-thought reasoning within the image itself? We introduce Pixel Reasoner, the first open-source framework that enables VLMs to “think in pixel space” through curiosity-driven reinforcement learning. Current VLMs reason only in text — even when grounded in rich images or videos, their logical steps are verbalized in natural language. This restricts their ability to interrogate visual evidence and demonstrate how conclusions are drawn. 🔍 So we ask: What if we could make VLMs "show their work" by reasoning directly in the pixel space? Inspired by GPT-o3’s "think-in-image" ability, we propose a framework where VLMs use interactive visual operations — zoom, select-frame, highlight — to reason through complex visual inputs. To do this, we design a two-stage training process: Instruction tuning with synthesized visual reasoning traces. Reinforcement learning with curiosity-driven reward to balance exploration between pixel and text reasoning ✨ With this, Pixel Reasoner achieves near-SoTA performance on many information-rich multimodal benchmarks: 📊 84% on InfographicsVQA 🧠 84% on V* benchmark 🧩 74% on TallyQA-Complex It also achieves strong accuracy of 68% on MVBench (a video benchmark). Website: Paper: Code: Demo: (coming soon)
Wenhu Chen82,829 次观看 • 1 年前

Tired of writing academic papers or articles, struggling to find relevant papers to cite? ScholarCopilot is here to help. It can help you autocomplete your writing, providing or suggesting relevant citations. Try out our demo at You can also set up your own demo with our Github code at
Wenhu Chen41,918 次观看 • 1 年前

Meet VideoScore, the FIRST fine-grained video reward/metric model! VideoScore can be used do rate synthesized videos or reward your T2V model through RLHF. We curate VideoFeedback, containing 37K human-annotated multi-aspect ratings for synthesized videos. VideoScore was trained on this dataset to simulate human judgement on five aspects including visual quality (VQ), temporal consistency (TC), dynamic degree (DD), text-alignment (TVA) and factualness (FC). VideoScore shows strong correlation with human raters. On four datasets: VideoFeedback-test, EvalCrafter (Ying Shan), VBench (Ziwei Liu) and GenAI-bench, VideoScore can universally beat other metric model by a huge margin. Notably, it outperforms "gpt-4o as judge" by 50% on our eval set in terms of spearman correlation. Arxiv: We release everything including data and model. Website: Demo: Work led by Xuan He and Dongfu Jiang. Lots of students contributed to the annotation pipeline. We also thank for providing compute to us.
Wenhu Chen14,585 次观看 • 2 年前
没有更多内容可加载