This model is a fine-tuned version of inflatebot/MN-12B-Mag-Mell-R1 on the kto_rp dataset. It achieves the following results on the evaluation set: — Loss: 0.3763 — Rewards/chosen: 0.6433 — Logps/chosen: -219.3819 — Logits/chosen: -923703389.9450 — Rewards/rejected: -1.3013 — Logps/rejected: -231.1543 — Logits/rejected: -831308124.5957 — Rewards/margins: 1.9446 — Kl: 0.0 The following hyperparameters were used during training: — learningrate: 5e-07 — trainbatchsize: 1 — evalbatchsize: 1 — seed: 42 — distributedtype: multi-GPU — numdevices: 4 — gradientaccumulations
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: taozi555
Теги: mistral, llama-factory, full, generated_from_trainer, conversational, text-generation-inference, endpoints_compatible
Лайков: 4 | Загрузок: 10
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.