A continued pretrained version of unsloth/Qwen2.5-7B model using unsloth’s low rank adaptation on a dataset of DTF posts. The adapter is already merged with the model. For pretraining, posts from SubMaroon/DTFcommentsResponsesCounts were selected, deduplicated by simple df.unique` and filtered by length of 1000 < x < 128000 tokens. The training dataset size was roughly 75M tokens. — NVidia Tesla A100 80GB: ~8.5 hours — NVidia RTX 3090ti: ~33.5 hours
Модальности:
Генерация текста
Задача: Генерация текста
Автор: chameleon-lizard
Теги: qwen2, ru
Лайков: 3 | Загрузок: 32
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.