Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled from a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia articles, and mathematical and coding contents, etc for continual pre-training. The instruction-tuned models (Instruct) were built by supervised fine-tuning (SFT) on the synthetic data specially built for Japanese (see the Training Datasets section for details). See the Swallow Model Index section to find other model variants. — November 11, 2024: Released Llama-3.1-Swallow-8B-v0.2 and Llama-3.1-Swallow-8B-Instruct-v0.2. — October 08, 2024: Released Llama-3.1-Swallow-8B-v0.1, Llama-3.1-Swallow-8B-Instruct-v0.1, Llama-3.1-Swallow-70B-v0.1, and Llama-3.1-Swallow-70B-Instruct-v0.1. The website https://swallow-llm.github.io/ provides large language models developed by the Swallow team. Model type: Please refer to Llama 3.1 MODELCARD for details on the model…
Модальности:
Генерация текста
Задача: Генерация текста
Автор: tokyotech-llm
Теги: llama, en, ja, text-generation-inference, endpoints_compatible
Лайков: 4 | Загрузок: 81
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.