— Model Architecture: Qwen2 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 — Activation quantization: FP8 — Intended Use Cases: Intended for commercial and research use in English. Similarly to Meta-Llama-3-8B-Instruct, this models is intended for assistant-like chat. — Out-of-scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. — Release Date: 6/14/2024 — Version: 1.0 — License(s): apache-2.0 — Model Developers: Neural Magic Quantized version of Qwen2-0.5B-Instruct. It achieves an average score of 42.94 on the OpenLLM benchmark (version 1), whereas the unquantized model achieves 42.96. This model was obtained by quantizing the weights and activations of Qwen2-0.5B-Instruct to FP8 data type, ready for inference with vLLM >= 0.5.0. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. Only the weights and activations of the linear operators within transformers blocks are quantized. Symmetric per-tensor quantization is applied, in which a single linear scaling maps the FP8…
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: RedHatAI
Теги: qwen2, fp8, vllm, conversational, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 1,761
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.