— Model Architecture: Llama-3.1-8B — Input: Text — Output: Text — Model Optimizations: — Sparsity: 2:4 — Weight quantization: FP8 — Activation quantization: FP8 — Release Date: 11/21/2024 — Version: 1.0 — License(s): llama3.1 — Model Developers: Neural Magic This is AI model especialized in grade-school math obtained by fine-tuning the 2:4 sparse Sparse-Llama-3.1-8B-2of4 on the GSM8k dataset, followed by one-shot quantization. It achieves 66.8% 0-shot accuracy on the test set of GSM8k, compared to 66.3% for the fine-tuned dense model Llama-3.1-8B-gsm8k — demonstrating over 100.0% accuracy recovery. In constrast, the pretrained Llama-3.1-8B achieves 50.7% 5-shot accuracy and the sparse foundational Sparse-Llama-3.1-8B-2of4 model achieves 56.3% 5-shot accuracy. This model was obtained by quantizing the weights of Sparse-Llama-3.1-8B-gsm8k-2of4 to INT4 data type. This optimization reduces the number of bits used to represent weights and activations from 16 to 8, reducing GPU memory requirements (by approximately 50%) and increasing matrix-multiply compute throughput (by approximately 2x). Weight quantization also reduces disk size requirements by approximately 50%. Only weights and…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: RedHatAI
Теги: llama, vllm, sparsity, quantized, conversational, en, compressed-tensors
Лайков: 3 | Загрузок: 32
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.