Magpie-Align/Llama-3-8B-Magpie-Align-v0.3 - Каталог нейросетей
Генерация текста

Magpie-Align/Llama-3-8B-Magpie-Align-v0.3

Добавлено:
Magpie-Align/Llama-3-8B-Magpie-Align-v0.3

Online Model Demo: https://huggingface.co/spaces/flydust/Chat-with-Magpie This model is an aligned version of meta-llama/Meta-Llama-3-8B. We apply the following pipeline: We first perform SFT using: Magpie-Align/Magpie-Pro-MT-300K-v0.1 Magpie-Align/Magpie-Reasoning-150K Magpie-Align/Magpie-Qwen2-Pro-200K-Chinese SFT Model Checkpoint: Magpie-Align/Llama-3-8B-Magpie-Align-SFT-v0.3 We then perform DPO on the princeton-nlp/llama3-ultrafeedback-armorm dataset. The overall performance is much better than the official Llama-3-8B-Instruct Model! Plus, it can answer Chinese queries frequently, thanks to our new Chinese instruction dataset! — Alpaca Eval 2 (vs GPT-4-Turbo-1106): 48.58 (LC), 50.36 (WR) — Alpaca Eval 2 (vs Llama-3-8B-Instruct): 73.65 (LC), 75.81 (WR) — Arena Hard: 42.2 — WildBench WB-Score: 41.1 — Zero-Eval GSM: 50.0 We compare our Llama-3-8B-Magpie-Align with official and other open-aligned LLMs that have been fine-tuned from base models and have publicly released their training datasets. The results are as follows: Conversation Template: Please use Llama 3 official chat template for the best performance. How to use it? Please check the official Llama 3 repository for…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: Magpie-Align
Теги: tensorboard, llama, alignment-handbook, axolotl, trl, dpo, sft, generated_from_trainer
Лайков: 3  |  Загрузок: 8,566

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.