QuantFactory/Qwen2.5-Lumen-14B-GGUF - Каталог нейросетей
Генерация текста

QuantFactory/Qwen2.5-Lumen-14B-GGUF

Добавлено:
QuantFactory/Qwen2.5-Lumen-14B-GGUF

This is quantized version of v000000/Qwen2.5-Lumen-14B created using llama.cpp Qwen direct preference optimization finetuned for ~3 epochs.* A qwen2.5 preference finetune, targeting prompt adherence, storywriting and roleplay. Trained Qwen2.5-14B-Instruct for 2 epochs on NVidia A100, and on dataset jondurbin/gutenberg-dpo-v0.1, saving different checkpoints along the way (completely different runs at varying epochs and learning rates). Tanliboy trained Qwen2.5-14B-Instruct for 1 epoch on HuggingFaceH4/ultrafeedbackbinarized, (Credit to Tanliboy! Check out the model here*) Mass checkpoint merged, Based on Qwen2.5-14B-Instruct (Base Model). Merged with a sophosympatheia’s SLERP gradient «Ultrafeedback-Binarized DPO» and «Gutenberg DPO»* Merged with a sophosympatheia’s SLERP gradient «Qwen2.5-14B-Instruct» and «Gutenberg DPO»* Merged all DPO checkpoints and SLERP variations with MODELSTOCK to analyze geometric properties and get the most performant aspects of all runs/merges. Model Stock was chosen due to the similarity between the merged models. This was chosen due to the fact that evaluation for ORPO* is unclear, so it’s hard to know which runs are the best. Temp 1.3 [1], MinP 0.012…

Модальности:
Генерация текста

Области применения:
Следование инструкциям Диалог / чат


Задача: Генерация текста
Автор: QuantFactory
Теги: gguf, qwen, qwen2.5, finetune, dpo, orpo, qwen2, chat
Лайков: 3  |  Загрузок: 498

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.