
Philip Schroeder
@philip_mit • 1,739 subscribers
PhD student at MIT in Computer Science. @MIT @MIT_CSAIL @MITEECS Robots, RL, embodied reasoning, video-language models.
Videos

Excited to introduce SOLE-R1, a video-language reasoning model for zero-shot reward prediction for robot manipulation tasks! SOLE-R1 reasoning can serve as the SOLE signal for learning new tasks (completely from scratch) through online RL - i.e., robots start with random actions and learn previously unseen tasks guided only by SOLE-R1 rewards, without any demonstrations, ground-truth rewards, success indicators, or task-specific tuning. SOLE-R1 significantly outperforms strong baselines (e.g., Robometer, RoboReward, TOPReward, GPT-5, Gemini-3-Pro) in zero-shot online RL when evaluated across 40 tasks - including a real-world tabletop manipulation setting and 4 sim environments (LIBERO, ManiSkill, Meta-World, RoboSuite). We open source all models, training data, and code. Website, demos, and paper at: 🧵 (1/6)
Philip Schroeder18,957 Aufrufe • vor 3 Monaten
Keine weiteren Inhalte verfügbar