This is a smaller version of the Meta’s Llama-3.2-1B decoder transformer model pretrained from scratch for 23 hours using a single A100 40GB GPU and 274 million tokens of Amharic text. — It has 400 Million parameters — The context size of this model is 1024 tokens. — It has the same tokenizer as Llama-3.2-1B, trained from scratch using the same Amharic dataset as the model with a vocabulary size of 32k. — Validation Perplexity: 41.3 — This is a base model and hasn’t undergone any supervised finetuing yet. First, you need to install the latest version of transformers You can use this model directly with a pipeline for text generation:
Модальности:
Генерация текста
Задача: Генерация текста
Автор: rasyosef
Теги: llama, am, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 68
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.