RedHatAI/Mixtral-8x7B-Instruct-v0.1-AutoFP8 - Каталог нейросетей
Генерация текста

RedHatAI/Mixtral-8x7B-Instruct-v0.1-AutoFP8

Добавлено:
RedHatAI/Mixtral-8x7B-Instruct-v0.1-AutoFP8

— Model Architecture: Mixtral-8x7B-Instruct-v0.1 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 — Activation quantization: FP8 — Intended Use Cases: Intended for commercial and research use in English. Similarly to Meta-Llama-3-7B-Instruct, this models is intended for assistant-like chat. — Out-of-scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. — Release Date: 6/8/2024 — Version: 1.0 — License(s): apache-2.0 — Model Developers: Neural Magic Quantized version of Mixtral-8x7B-Instruct-v0.1. It achieves an average score of 73.19 on the OpenLLM benchmark (version 1), whereas the unquantized model achieves 73.48. This model was obtained by quantizing the weights and activations of Mixtral-8x7B-Instruct-v0.1 to FP8 data type, ready for inference with vLLM >= 0.5.0. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. Only the weights and activations of the linear operators within transformers blocks are quantized. Symmetric per-tensor quantization is applied, in which a…

Модальности:
Генерация текста

Области применения:
Диалог / чат Следование инструкциям


Задача: Генерация текста
Автор: RedHatAI
Теги: mixtral, fp8, vllm, conversational, text-generation-inference, endpoints_compatible
Лайков: 3  |  Загрузок: 38

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.