Doge uses wsdscheduler as the training scheduler, which divides the learning rate into three stages: warmup, stable, and decay. It allows us to continue training on any new dataset from any checkpoint in the stable stage` without spikes of the training. Here are the initial learning rates required to continue training at each checkpoint: — Doge-20M: 8e-3 — Doge-60M: 6e-3 — Doge-160M: 4e-3 — Doge-320M: 2e-3
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: SmallDoge
Теги: doge, conversational, custom_code, en, endpoints_compatible
Лайков: 3 | Загрузок: 79
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.