MobileLLM is introduced: «MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases», published in ICML 2024. Model Architecture: MobileLLM is an auto-regressive language model leveraging an optimized transformer architecture, specifically engineered for on-device applications with constrained resources. MobileLLM integrated several key techniques including: (1) SwiGLU activation function, (2) deep and thin architectures, (3) embedding sharing, (4) grouped-query attention. MobileLLM-125M/350M attains a remarkable 2.7%/4.3% accuracy boost over preceding 125M/350M SoTA models on zero-shot commonsense reasoning tasks. In our updated version, we further demonstrate that our design philosophy scales effectively to larger models, with SoTA results for MobileLLM-600M/1B/1.5B. To load the pretrained model for further finetuning or evaluation: Note that the default tokenizer does not contain special tokens. For example you can use: We provide the pretraining code in https://github.com/facebookresearch/MobileLLM We also provide evaluation script for calculating ppl of wikitext-2 test split: It takes the following number of days to train MobileLLM on 1T tokens using…
Модальности:
Генерация текста
Задача: Генерация текста
Автор: vonjack
Теги: llama, model-index, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 94
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.