mlabonne/phixtral-3x2_8 - Каталог нейросетей
Генерация текста

mlabonne/phixtral-3x2_8

Добавлено:
mlabonne/phixtral-3x2_8

phixtral-3x2_8 is the first Mixure of Experts (MoE) made with two microsoft/phi-2 models, inspired by the mistralai/Mixtral-8x7B-v0.1 architecture. It performs better than each individual expert. The evaluation was performed using LLM AutoEval on Nous suite. Check YALL — Yet Another LLM Leaderboard to compare it with other models. The model has been made with a custom version of the mergekit library (mixtral branch) and the following configuration: Here’s a Colab notebook to run Phixtral in 4-bit precision on a free T4 GPU. Inspired by mistralai/Mixtral-8x7B-v0.1, you can specify the numexpertspertok and numlocalexperts in the config.json file (2 for both by default). This configuration is automatically loaded in configuration.py`. vince62s implemented the MoE inference code in the modelingphi.py` file. In particular, see the MoE class. A special thanks to vince62s for the inference code and the dynamic configuration of the number of experts. He was very patient and helped me to debug everything. Thanks to Charles Goddard for the mergekit library and the implementation of the MoE for clowns. Thanks to ehartford and lxuechen for their fine-tuned phi-2 models.

Модальности:
Генерация текста

Области применения:
Генерация кода Диалог / чат


Задача: Генерация текста
Автор: mlabonne
Теги: phi-msft, moe, nlp, code, cognitivecomputations/dolphin-2_6-phi-2, lxuechen/phi-2-dpo, conversational, custom_code
Лайков: 3  |  Загрузок: 35

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.