This is a llamafied version of bagel-dpo-20b-v04, which is a fine-tune of internlm2-20b, which underwent additional fine-tuning using direct preference optimization (DPO). The non-DPO version is available here, and is likely superior for roleplay. Compute for the SFT phase was generously provided by MassedCompute Compute for the DPO phase was generously provided by latitude.sh There are many data sources used in the bagel models. See https://github.com/jondurbin/bagel for more information. Only train splits are used, and a decontamination by cosine similarity is performed at the end as a sanity check against common benchmarks. If you don’t know the difference between train and test, please learn. — ai2arc — Abstraction and reasoning dataset, useful in measuring «intelligence» to a certain extent. — airoboros — Variety of categories of synthetic instructions generated by gpt-4. — apps — Python coding dataset with 10k problems. — belebele — Multi-lingual reading comprehension dataset. — bluemoon — Roleplay data scraped from Bluemoon, then cleaned and formatted as ShareGPT. — boolq — Corpus of yes/no questions (which can be surprisingly difficult for AI to answer apparently?) -…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: jondurbin
Теги: llama, conversational, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 18
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.