RedHatAI/Sparse-Llama-3.1-8B-ultrachat_200k-2of4-quantized.w4a16 - Каталог нейросетей
Генерация текста

RedHatAI/Sparse-Llama-3.1-8B-ultrachat_200k-2of4-quantized.w4a16

Добавлено:
RedHatAI/Sparse-Llama-3.1-8B-ultrachat_200k-2of4-quantized.w4a16

— Model Architecture: Llama-3.1-8B — Input: Text — Output: Text — Model Optimizations: — Sparsity: 2:4 — Weight quantization: INT4 — Release Date: 11/21/2024 — Version: 1.0 — License(s): llama3.1 — Model Developers: Neural Magic This is a multi-turn conversational AI model obtained by fine-tuning the 2:4 sparse Sparse-Llama-3.1-8B-2of4 on the ultrachat200k dataset, followed by quantization. On the AlpacaEval benchmark (version 1), it achieves a score of 61.6, compared to 62.0 for the fine-tuned dense model Llama-3.1-8B-ultrachat200k — demonstrating a 99.4% accuracy recovery. This model was obtained by quantizing the weights of Sparse-Llama-3.1-8B-ultrachat200k-2of4 to INT4 data type. This optimization reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 75%. That is on top of the reduction of 50% of weights via 2:4 pruning employed on Sparse-Llama-3.1-8B-ultrachat200k-2of4. Only the weights of the linear operators within transformers blocks are quantized. Symmetric per-channel quantization is applied, in which a linear scaling per output dimension maps the INT4 and floating point representations of the quantized…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: RedHatAI
Теги: llama, vllm, sparsity, conversational, en, compressed-tensors
Лайков: 3  |  Загрузок: 38

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.