This is a finetune of Qwen2.5-Instruct on WiroAI/dolphin-r1-french. — DeepSeek’s distilled models sometimes reason in Chinese or English even though prompted in another language. — Open-Source models still need improvement on relatively low-resource languages. — A motivation to reproduce R1 and contribute to the community. — We train the model on the WiroAI/dolphin-r1-french for 2 epochs. We use learning rate of 1e-5 and max seq length 4096. The training follows a cosine learning rate schedule with a 10% warmup phase. — Training took 5 days in 8xA6000 ADA cluster. — Normally, R1 team compares the performance of OpenR1 models to DeepSeek-Distill-Qwen-7B and OpenThinker-7B using lighteval. However, the datasets are only MATH oriented so not to conclude anything we won’t disclose the default results. You can find the training and evaluation code at: https://github.com/huggingface/open-r1/ — We observed that reasoning process has slightly improved. Our model thinks more clearly in French compared to the DeepSeek’s reasoning model. — This model trained for experimental motives and any benchmark evaluation is appreciated. Please be aware that this model will be producing more tokens…
Модальности:
Генерация текста
Области применения:
Логика и рассуждение Диалог / чат
Задача: Генерация текста
Автор: WiroAI
Теги: qwen2, generated_from_trainer, trl, sft, reasoning, thinking, deepseek, dolphin
Лайков: 4 | Загрузок: 16
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.