lamm-mit/PRefLexOR_ORPO_DPO_EXO_REFLECT_10222024 - Каталог нейросетей
Генерация текста

lamm-mit/PRefLexOR_ORPO_DPO_EXO_REFLECT_10222024

Добавлено:
lamm-mit/PRefLexOR_ORPO_DPO_EXO_REFLECT_10222024

We introduce PRefLexOR (Preference-based Recursive Language Modeling for Exploratory Optimization of Reasoning), a framework that combines preference optimization with concepts from Reinforcement Learning (RL) to enable models to self-teach through iterative reasoning improvements. Central to PRefLexOR are thinking tokens, which explicitly mark reflective reasoning phases within model outputs, allowing the model to recursively engage in multi-step reasoning, revisiting, and refining intermediate steps before producing a final output. The foundation of PRefLexOR lies in Odds Ratio Preference Optimization (ORPO), where the model learns to align its reasoning with human-preferred decision paths by optimizing the log odds between preferred and non-preferred responses. The integration of Direct Preference Optimization (DPO) further enhances model performance by using rejection sampling to fine-tune reasoning quality, ensuring nuanced preference alignment. This hybrid approach between ORPO and DPO mirrors key aspects of RL, where the model is continuously guided by feedback to improve decision-making and reasoning. Active learning mechanisms allow PRefLexOR to dynamically generate new…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: lamm-mit
Теги: llama, PRefLexOR, conversational, text-generation-inference, endpoints_compatible
Лайков: 3  |  Загрузок: 20

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.