This continual pretrained model is pretrained on roughly 2.4 Billion tokens of Dutch language data based on Wikipedia and MC4. Primary objective with continual pretraining on Dutch was to make the model more ‘fluent’ when using the Dutch language. It will also have gained some additional Dutch knowledge. As a base model the IBM Granite 3.0 2B Instruct model was used. See ibm-granite/granite-3.0-2b-instruct for all information about the IBM Granite foundation model. A basic example of how to use this continual pretrained model. !! IMPORTANT NOTE !! As this is an instruct model that was continual pretrained on dutch data there is some degredation in the performance regarding instruction-following. This custom pretrained model should be further finetuned with SFT in which the embedding and lm_head layer are also trained. Given a proper SFT dataset in dutch this will restore the instruction following/EOS token functionality. See the SFT training notebook for Schaapje on one of the ways on how to do this. As with all LLM’s this model can also experience bias and hallucinations. Regardless of how you use this model always perform the necessary testing and validation. The datasets used…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: robinsmits
Теги: tensorboard, granite, granite 3.0, schaapje, conversational, nl
Лайков: 3 | Загрузок: 37
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.