— Model creator: Nexusflow — Original model: Starling-LM-7B-beta — Developed by: Banghua Zhu , Evan Frick , Tianhao Wu , Hanlin Zhu, Karthik Ganesan, Wei-Lin Chiang, Jian Zhang, and Jiantao Jiao. — Model type: Language Model finetuned with RLHF / RLAIF — License: Apache-2.0 license under the condition that the model is not used to compete with OpenAI — Finetuned from model:** Openchat-3.5-0106 (based on Mistral-7B-v0.1) We introduce Starling-LM-7B-beta, an open large language model (LLM) trained by Reinforcement Learning from AI Feedback (RLAIF). Starling-LM-7B-beta is trained from Openchat-3.5-0106 with our new reward model Nexusflow/Starling-RM-34B and policy optimization method Fine-Tuning Language Models from Human Preferences (PPO). Harnessing the power of our ranking dataset, berkeley-nest/Nectar, our upgraded reward model, Starling-RM-34B, and our new reward training and policy tuning pipeline, Starling-LM-7B-beta scores an improved 8.12 in MT Bench with GPT-4 as a judge. Stay tuned for our forthcoming code and paper, which will provide more details on the whole process. AWQ is an efficient, accurate and blazing-fast low-bit weight quantization method, currently…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Языки программирования:
Rust
Задача: Генерация текста
Автор: solidrust
Теги: mistral, reward model, RLHF, RLAIF, quantized, 4-bit, AWQ, autotrain_compatible
Лайков: 3 | Загрузок: 22
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.