This model is a fine-tuned version of mistralai/Mistral-7B-v0.1 on the Rijgersberg/norobotsnl and Rijgersberg/ultrachat10knl datasets. It achieves the following results on the evaluation set: — Loss: 1.0263 In order to investigate the effect of pretraining Rijgersberg/GEITje-7B on the finetuning of Rijgersberg/GEITje-7B-chat, I also subjected the base model Mistral 7B v0.1 to the exact same training. This model is called Mistral-7B-v0.1-chat-nl. Read more about GEITje and GEITje-chat in the 📄 README on GitHub. The following hyperparameters were used during training: — learningrate: 1e-05 — trainbatchsize: 2 — evalbatchsize: 8 — seed: 42 — gradientaccumulationsteps: 8 — totaltrainbatchsize: 16 — optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08 — lrschedulertype: cosine — lrschedulerwarmupratio: 0.1 — numepochs: 3 — Transformers 4.36.0.dev0 — Pytorch 2.1.1+cu121 — Datasets 2.15.0 — Tokenizers 0.15.0
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: Rijgersberg
Теги: tensorboard, mistral, generated_from_trainer, GEITje, conversational, nl, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 24
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.