i wanted to learn more about exposure bias mitigation in language models and came across ReMask. it’s a neat idea, and i wanted to give it a go. — during training, the model processes input sequences twice — once with the full sequence & once with masked sequence. — computes model outputs for both. — divergence loss is computed as the average of forward and backward KL divergences. — final loss is a weighted sum of the cross entropy losses and the divergence loss.
Модальности:
Генерация текста
Задача: Генерация текста
Автор: aloobun
Теги: llama, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 25
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.