This is a merge of pre-trained language models created using mergekit. I doubt this lass is as well-poisoned as her Nemo counterpart; Qwen 2.5 presumably uses enough synthetic data that she’s not as inclined to shake her natural alignment and tendencies. (Additionally I did have a higher proportion of nontoxic data in the DPO, and trained based on a non-Unleashed model. Then undid a lot of my work with the SLERP, technically. 😉 ) Edit: Tagging with not for all audiences because of the dataset, I guess, but she’s actually still aligned? Not sure what niche she falls into exactly. But that doesn’t matter because she got 78.0391 in EQ-Bench and 100% parseable. 😀 Clearly worthy in her own right. In addition to basing the model on a variation of Lumen that I attempted to rebase and heal from the damage of original instruct, I did my own DPO including these datasets: Then used the SLERP (sophosympatheia gradient) to heal some of the damage that inevitably wrought on the intermediate layers with the model before training. The following models were included in the merge: Lambent/arsenic-v0.1-dpo-qwen2.5-14B Lambent/qwen2.5-reinstruct-alternate-lumen-14B The following YAML…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: Lambent
Теги: qwen2, mergekit, merge, not-for-all-audiences, conversational, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 18
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.