Harley-ml/Dillionv2-1.3M - Каталог нейросетей
Генерация текста

Harley-ml/Dillionv2-1.3M

Добавлено:
Harley-ml/Dillionv2-1.3M

Dillionv2 is our second generation model of the Dillion SLM family. It is a significant improvement over v1 (in everything except ARC). We trained Dillionv2 for one epoch on 24B tokens for a combined total of 35 hours on an RTX 2060 and two T4s from Kaggle with a batch size of 384 and a gradient accumulation of 2. The dataset is 34B tokens (we only use the first 24B) and 146GB in total: 1. FineWeb-edu (35GB): Educational-filtered Common Crawl 2. DCLM-Edu (20GB): Educational-filtered webtext 3. The Pile Deduped (20GB): Broad, diverse 23-source dataset 4. FineWeb-HQ (20GB): Knowledge-filtered Webtext 5. FineMath (13GB): Math-filtered Common Crawl 6. Cosmopedia-v2 (7GB): Synthetic textbooks 7. Wikipedia (5GB): you better know what this is 8. NpSetPython-Edu (3.5GB): normalized Python code 9. Misc (600MB): LessWrong + HF configs + HF dataset/model cards The final loss ended at 3.078, which is a perplexity of 21.417. Dillionv2 shows stonger performace on multiple benchmarks than v1, except ARC. For a comphrehensive comparison among many small models, including my own, such as this one, go to AxiomicLab’s Open SLM Leaderboard. 1. Educational research, learning, etc 2. fine-tuning for…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: Harley-ml
Теги: qwen3_5_text, dillion, small, dillionv2, Harley-ml, slm, en
Лайков: 4  |  Загрузок: 456

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.