 This is quantized version of princeton-nlp/gemma-2-9b-it-DPO created using llama.cpp This model was trained under the same setup as gemma-2-9b-it-SimPO, with the DPO objective. SimPO (Simple Preference Optimization) is an offline preference optimization algorithm designed to enhance the training of large language models (LLMs) with preference optimization datasets. SimPO aligns the reward function with the generation likelihood, eliminating the need for a reference model and incorporating a target reward margin to boost performance. Please refer to our preprint and github repo for more details. We fine-tuned google/gemma-2-9b-it on princeton-nlp/gemma2-ultrafeedback-armorm with the DPO objective. — Developed by: Yu Meng, Mengzhou Xia, Danqi Chen — Model type: Causal Language Model — License: gemma — Finetuned from model: google/gemma-2-9b-it — Repository: https://github.com/princeton-nlp/SimPO — Paper: https://arxiv.org/pdf/2405.14734 We use princeton-nlp/gemma2-ultrafeedback-armorm as the…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: QuantFactory
Теги: gguf, alignment-handbook, generated_from_trainer, endpoints_compatible, conversational
Лайков: 3 | Загрузок: 1,332
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.