Pyroserenus/L3-8B-Stheno-v3.3-32K-8.0bpw-h8-exl2 - Каталог нейросетей
Генерация текста

Pyroserenus/L3-8B-Stheno-v3.3-32K-8.0bpw-h8-exl2

Добавлено:
Pyroserenus/L3-8B-Stheno-v3.3-32K-8.0bpw-h8-exl2

Trained with compute from Backyard.ai | Thanks to them and @dynafire for helping me out. Relevant Axolotl Configurations: -> Taken from winglian/Llama-3-8b-64k-PoSE — I tried to find my own configs, hours of tinkering but the one he used worked best, so I stuck to it. — 2M Rope Theta had the best loss results during training compared to other values. — Leaving it at 500K rope wasn’t that much worse, but 4M and 8M Theta made the gradnorm values worsen even if loss drops fast. — Mixing in Pretraining Data was a PITA. Made it a lot worse with formatting. — Pretraining / Noise made it worse at Haystack too? It wasn’t all Green, Mainly Oranges. — Improper / Bad Rope Theta shows in GradNorm exploding to thousands. It’ll drop to low values alright, but it’s a scary fast drop even with gradient clipping.

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: Pyroserenus
Теги: llama, conversational, en, text-generation-inference, endpoints_compatible, 8-bit, exl2
Лайков: 3  |  Загрузок: 20

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.