HenryHHHH/DistilLlama - Каталог нейросетей
Генерация текста

HenryHHHH/DistilLlama

Добавлено:
HenryHHHH/DistilLlama

This model is a distilled version of LLaMA 2, containing approximately 80 million parameters. It was trained using a mix of OpenWebText and WikiText Raw V1 datasets. Knowledge distillation was employed to transfer knowledge from a larger «teacher» model—Meta’s 7B LLaMA 2—to help this smaller model mimic the behavior of the teacher. The architecture is based on LLaMA 2, with the following parameters: During each training step, the input data ( X ) is fed to both the teacher and student models. The student model calculates output logits and loss with the true labels, while the teacher model only generates logits. The total loss combines task-specific loss and distillation loss: — Batch Size: 64 — Max Sequence Length: 128 — Epochs: 2 — Log Interval: 3000 — Learning Rate: 3e-4 — Warmup Steps: 4000 — Accumulation Steps: 8 — Load Model: True — Temperature: 2.0 — Alpha: 0.3 The model’s performance is evaluated on 200 queries created in-house. For more details, visit the GitHub repository. 1. Input: The capital of France is — Output: «The capital of France is located in the southern province of Lyon, France. The capital is the main hub of the French capital, La Caillion, and the main…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: HenryHHHH
Теги: llama, knowledge-distillation, causal-lm, openwebtext, wikitext, transfer-learning, en, text-generation-inference
Лайков: 3  |  Загрузок: 84

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.