Qwen2.5-1.5B-Instruct-QwQ is a fine-tuned model based on Qwen2.5-1.5B-Instruct. It was fine-tuned on roughly 20k samples from QwQ-32B-Preview. Compared to Qwen2.5-1.5B-Instruct, this fine-tuned model seems more performant in mathematics contexts and general reasoning. Also it shows some capabilities of self-correction, altough it seems a bit limited (bigger models seem to learn self-correction better, e.g. the 3B & 7B version show much better self-correction abilities in my experiments). For data generation, math problems from the train sets of the GSM8k and MATH datasets were used. This repo contains the instruction-tuned 1.5B Qwen2.5 model fine-tuned on QwQ reasoning chains, which has the following features: — Type: Causal Language Models — Training Stage: Pretraining & Post-training — Architecture: transformers with RoPE, SwiGLU, RMSNorm, Attention QKV bias and tied word embeddings — Number of Parameters: 1.54B — Number of Paramaters (Non-Embedding): 1.31B — Number of Layers: 28 — Number of Attention Heads (GQA): 12 for Q and 2 for KV — Context Length
Модальности:
Генерация текста
Области применения:
Математика Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: micaebe
Теги: qwen2, chat, trl, sft, math, conversational, en, model-index
Лайков: 4 | Загрузок: 50
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.