LoneStriker/Starling-LM-7B-beta-8.0bpw-h8-exl2 - Каталог нейросетей
Генерация текста

LoneStriker/Starling-LM-7B-beta-8.0bpw-h8-exl2

Добавлено:
LoneStriker/Starling-LM-7B-beta-8.0bpw-h8-exl2

— Developed by: Banghua Zhu , Evan Frick , Tianhao Wu , Hanlin Zhu, Karthik Ganesan, Wei-Lin Chiang, Jian Zhang, and Jiantao Jiao. — Model type: Language Model finetuned with RLHF / RLAIF — License: Apache-2.0 license under the condition that the model is not used to compete with OpenAI — Finetuned from model:** Openchat-3.5-0106 (based on Mistral-7B-v0.1) We introduce Starling-LM-7B-beta, an open large language model (LLM) trained by Reinforcement Learning from AI Feedback (RLAIF). Starling-LM-7B-beta is trained from Openchat-3.5-0106 with our new reward model Nexusflow/Starling-RM-34B and policy optimization method Fine-Tuning Language Models from Human Preferences (PPO). Harnessing the power of our ranking dataset, berkeley-nest/Nectar, our upgraded reward model, Starling-RM-34B, and our new reward training and policy tuning pipeline, Starling-LM-7B-beta scores an improved 8.12 in MT Bench with GPT-4 as a judge. Stay tuned for our forthcoming code and paper, which will provide more details on the whole process. For more detailed discussions, please check out our original blog post, and stay tuned for our upcoming code and paper! — Blog: https://starling.cs.berkeley.edu/ -…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: LoneStriker
Теги: mistral, reward model, RLHF, RLAIF, conversational, en, text-generation-inference, endpoints_compatible
Лайков: 3  |  Загрузок: 17

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.