OPRMs (Outcome & Process Reward Models) are trained to predict the correctness of each step on the position of «nn», as well as the correctness of the whole solution on the position of «».
Модальности:
Генерация текста
Области применения:
Математика
Задача: Генерация текста
Автор: ScalableMath
Теги: llama, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 10
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.