EmertonOmniBeagle-7B-dpo is a DPO fine-tune of mlabonne/Monarch-7B using the yleo/emertondpopairsjudge preference dataset created from Intel/orcadpo_pairs by replacing gpt 3.5 answer by a gpt4 Turbo answer. Then, LLM-Blender is used to judge between GPT4 and GPT4 Turbo. This model uses a context window of 8k. It is compatible with different templates, like chatml and Llama’s chat template.
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: yleo
Теги: mistral, dpo, conversational, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 30
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.