Instead of training an additional reward model that is likely to be gamed, we directly train the model on the social games! 🕹️ 🎲 🎮 This is the second step of Stable Alignment project, which is a supervised fine-tuned model on Anthropic HH-RLHF dataset (only on the ‘accepted’ options). Although this project aims to better align current LMs with social norms, inappropriate content and inherent biases in the training data will still impair the alignment of the model. The model should not be used directly in any application, without a prior assessment of safety and fairness concerns specific to the application. Please cite our paper if you use the data or code in this repo:
Модальности:
Генерация текста
Задача: Генерация текста
Автор: agi-css
Теги: llama, rlhf, alignment, simulation, computational social science, en, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 62
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.