Trained on a flavorful melange of the WizardLM, Airoboros, and Wizard Vicuna datasets. This model was trained using both linear and NTK-aware RoPE scaling in tandem. When loading, ensure that compressposemb (or scale) is set to 2, and alphavalue is set to 4. Both* values must be set. Expect context length of up to 8192 to work for sure. It will probably maintain coherence into the ~12k range, but I have not tested that.
Модальности:
Генерация текста
Задача: Генерация текста
Автор: chargoddard
Теги: llama, custom_code, en, endpoints_compatible
Лайков: 3 | Загрузок: 7
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.