firefly-qwen1.5-en-7b and firefly-qwen1.5-en-7b-dpo-v0.1 are trained based on Qwen1.5-7B to act as a helpful and harmless AI assistant. We use Firefly to train our models on a single V100 GPU with QLoRA. firefly-qwen1.5-en-7b is fine-tuned based on Qwen1.5-7B with English instruction data, and firefly-qwen1.5-en-7b-dpo-v0.1 is trained with Direct Preference Optimization (DPO) based on firefly-qwen1.5-en-7b. Our models outperform official Qwen1.5-7B-Chat, Gemma-7B-it, Zephyr-7B-Beta on Open LLM Leaderboard. Although our models are trained with English data, you can also try to chat with models in Chinese because Qwen1.5 is also good at Chinese. But we have not evaluated the performance in Chinese yet. We evaluate our models on Open LLM Leaderboard, they achieve good performance. The chat templates of our chat models are the same as Official Qwen1.5-7B-Chat: Both in SFT and DPO stages, We only use a single V100 GPU with QLoRA, and we use Firefly to train our models. The following hyperparameters are used during SFT: — numepochs: 1 — learningrate: 2e-4 — totaltrainbatchsize: 32 — maxseqlength: 2048 — optimizer: pagedadamw32bit — lrschedulertype: constantwithwarmup — warmupsteps: 700…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: YeungNLP
Теги: qwen2, conversational, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 60
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.