The Aquila-135M model is a small bilingual(Chinese and English) language model, which is trained using a two-phrase paradigm: pre-training and annealing. This model used 1.66TB bilingual tokens in Chinese and English during pre-training phrase and 100B tokens during annealing training phrase. In annealing stage, we selected 100B tokens of high-quality bilingual data and finally got our model. The Aquila-135M-Instuct model is finetuned using Infinity Instruct. The entire training process was conducted using FlagGems based on Triton and parallel training framework named FlagScale. — 2024/12/24: We have released Aquila-135M and Aquila-135M-Instruct. — 2024/12/24: We have released all datasets and intermediate checkpoints during training. Please feel free to use these models for analysis and experimentation. We have open-sourced all bilingual datasets during both pre-training and annealing phrases. Datasets composition and mix proportions are shown in the figure below. We followed the same evaluation setting of SmolLM models and evaluated models using the lighteval tool. The parameter count excludes the embedding part and Aquila-135M and SmolLM2-135M share an identical model…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: BAAI
Теги: mistral, conversational, en, zh, text-generation-inference, endpoints_compatible
Лайков: 4 | Загрузок: 2
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.