RewardAnything: Generalizable Principle-Following Reward Models Zhuohao Yu1,§ Jiali Zeng2 Weizheng Gu1 Yidong Wang1 Jindong Wang3 Fandong Meng2 Jie Zhou2 Yue Zhang4 Shikun Zhang1 Wei Ye1,† 1Peking University 2WeChat AI 3William & Mary 4Westlake University §Work done during Zhuohao’s internship at Pattern Recognition Center, WeChat AI, Tencent Inc; †Corresponding author. Traditional reward models learn implicit preferences from fixed datasets, leading to static judgments that struggle with the nuanced and multifaceted nature of human values. We believe that, much like Large Language Models follow diverse instructions, reward models must be able to understand and follow explicitly specified principles. RewardAnything embodies this new paradigm. Our models are designed to interpret natural language principles at inference time, enabling dynamic adaptation to a wide array of evaluation criteria without costly retraining. This approach shifts from fitting a single preference distribution to achieving true principle-following generalization. — 🧠 Principle-Following: Directly interprets and applies reward…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: zhuohaoyu
Теги: qwen3, reward-model, rlhf, principle-following, qwen, conversational, en, zh
Лайков: 4 | Загрузок: 13
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.