Загрузка видео...

Не удалось загрузить видео

На главную

Personalization assumes you need history with a user. What if you don't? Cold-start is hard: each task&user has many preference dimensions, but each user only cares about a few. A few strategic questions is all you need, if u know how preferences correlate across population👉🏻🧵

32,118 просмотров • 7 месяцев назад •via X (Twitter)

Комментарии: 12

Фото профиля Stella Li
Stella Li7 месяцев назад

Multi-turn RL seems appealing—interact w user, get a reward from the final response But RL collapses to asking everyone the SAME thing On personalized AIME, RL-trained policy asks identical Qs regardless of user answers. It learns to ignore you. 0% adaptivity🚨

Фото профиля Stella Li
Stella Li7 месяцев назад

Why? Complete user preference profiles have rich per-criterion structure of how preferences correlate across users. RL collapses all that into one terminal scalar and has to reverse-engineer it through credit assignment. The structure was there all along. RL just threw it away.

Фото профиля Stella Li
Stella Li7 месяцев назад

We propose✨PEP✨(Preference Elicitation with Priors): Instead of sparse rewards—Exploit the structure directly🧠 [1] Offline: learn preference correlations from complete profiles🔄 [2] Online: Bayesian inference to updates beliefs about all dimensions. No retraining needed.

Фото профиля Stella Li
Stella Li7 месяцев назад

How many questions does it take to truly understand someone? ➡️1-2 with PEP matches what RL needs 7-15 for. On social/commonsense tasks, RL never catches PEP's adaptive single-question performance—even after 15 turns. Because the world model infers what you didn't ask🔮

Фото профиля Stella Li
Stella Li7 месяцев назад

Across medical, math, social & commonsense reasoning: 📈80.8% preference alignment vs 68.5% for RL 🔀Adapts follow-up questions 39–62% of the time vs 0–28% for RL ⚙️~10K parameters vs 8B The bottleneck of preference elicitation is⭐️preference structure⭐️ used for reasoning.

Фото профиля Stella Li
Stella Li7 месяцев назад

Ablation confirms it: Remove preference correlations → flatlines at population average (similar to what RL converges to) no matter how many questions are asked📉 The world model is the core driver. Adaptive questioning saves one extra question on top of that.

Фото профиля Stella Li
Stella Li7 месяцев назад

The bottleneck in cold-start personalization is HOW to use the🤏data. Population preference profiles contain rich correlation structure—PEP exploits it directly via Bayesian inference⭐️ Maybe RL isn't the solution to every problem. Sometimes the right inductive bias beats scale

Фото профиля Stella Li
Stella Li7 месяцев назад

📄 💻 Co-lead with the amazing @avibose22 (who's on the job market🌟!!) Huge thanks to our mentors @faeze_brh @PangWeiKoh @SimonShaoleiDu @maryamfaz5050 @tsvetshop Lin Xiao @real_asli for making this project possible!🩵

Фото профиля Erik Wang
Erik Wang7 месяцев назад

I’ve never seen a better short-form medium for communicating research than this video 😆

Фото профиля Shobhnik
Shobhnik7 месяцев назад

which tool did you use to make this video/song😭

Фото профиля Michel aka Agent B
Michel aka Agent B7 месяцев назад

Very interesting work Stella. Thanks for the pointer 🙏

Фото профиля cannedapricots
cannedapricots7 месяцев назад

yokay

Похожие видео