vonjack/MobileLLM-125M-HF - Каталог нейросетей
Генерация текста

vonjack/MobileLLM-125M-HF

Добавлено:
vonjack/MobileLLM-125M-HF

MobileLLM is introduced: «MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases», published in ICML 2024. Model Architecture: MobileLLM is an auto-regressive language model leveraging an optimized transformer architecture, specifically engineered for on-device applications with constrained resources. MobileLLM integrated several key techniques including: (1) SwiGLU activation function, (2) deep and thin architectures, (3) embedding sharing, (4) grouped-query attention. MobileLLM-125M/350M attains a remarkable 2.7%/4.3% accuracy boost over preceding 125M/350M SoTA models on zero-shot commonsense reasoning tasks. In our updated version, we further demonstrate that our design philosophy scales effectively to larger models, with SoTA results for MobileLLM-600M/1B/1.5B. To load the pretrained model for further finetuning or evaluation: Note that the default tokenizer does not contain special tokens. For example you can use: We provide the pretraining code in https://github.com/facebookresearch/MobileLLM We also provide evaluation script for calculating ppl of wikitext-2 test split: It takes the following number of days to train MobileLLM on 1T tokens using…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: vonjack
Теги: llama, model-index, text-generation-inference, endpoints_compatible
Лайков: 3  |  Загрузок: 94

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.