— Model Architecture: Meta-Llama-3.2 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 — Activation quantization: FP8 — Intended Use Cases: Intended for commercial and research use in multiple languages. Similarly to Llama-3.2-3B-Instruct, this models is intended for assistant-like chat. — Out-of-scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. — Release Date: 9/25/2024 — Version: 1.0 — License(s): llama3.2 — Model Developers: Neural Magic Quantized version of Llama-3.2-3B-Instruct. It achieves an average score of 50.88 on a subset of task from the OpenLLM benchmark (version 1), whereas the unquantized model achieves 51.70. This model was obtained by quantizing the weights and activations of Llama-3.2-3B-Instruct to FP8 data type, ready for inference with vLLM built from source. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. Only the weights and activations of the linear operators within transformers blocks are quantized. Symmetric per-channel quantization is…
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: RedHatAI
Теги: llama, fp8, vllm, conversational, en, de, fr, it
Лайков: 3 | Загрузок: 504,775
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.