— Model Architecture: Mixtral-8x22B-Instruct-v0.1 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 — Activation quantization: FP8 — Intended Use Cases: Intended for commercial and research use in English. Similarly to Meta-Llama-3-8B-Instruct, this models is intended for assistant-like chat. — Out-of-scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. — Release Date: 8/11/2024 — Version: 1.1 — License(s): apache-2.0 — Model Developers: Neural Magic Quantized version of Mixtral-8x22B-Instruct-v0.1 with the updated tokenizer. It achieves an average score of 79.04 on the OpenLLM benchmark (version 1), whereas the unquantized model achieves 79.93. This model was obtained by quantizing the weights and activations of Mixtral-8x22B-Instruct-v0.1 to FP8 data type, ready for inference with vLLM >= 0.5.0. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. Only the weights and activations of the linear operators within transformers blocks are quantized. Symmetric per-tensor…
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: RedHatAI
Теги: mixtral, fp8, vllm, conversational, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 37
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.