RedHatAI/Meta-Llama-3.1-70B-FP8 - Каталог нейросетей
Генерация текста

RedHatAI/Meta-Llama-3.1-70B-FP8

Добавлено:
RedHatAI/Meta-Llama-3.1-70B-FP8

— Model Architecture: Meta-Llama-3.1 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 — Activation quantization: FP8 — Intended Use Cases: Intended for commercial and research use in multiple languages. Similarly to Meta-Llama-3.1-8B, this model serves as a base version. — Out-of-scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. — Release Date: 7/23/2024 — Version: 1.0 — License(s): llama3.1 — Model Developers: Neural Magic Quantized version of Meta-Llama-3.1-70B. It achieves an average score of 79.70 on the OpenLLM benchmark (version 1), whereas the unquantized model achieves 79.84. This model was obtained by quantizing the weights and activations of Meta-Llama-3.1-70B to FP8 data type, ready for inference with vLLM built from source. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. Only the weights and activations of the linear operators within transformers blocks are quantized. Symmetric per-tensor quantization is applied, in which a single linear scaling maps…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: RedHatAI
Теги: llama, fp8, vllm, en, de, fr, it, pt
Лайков: 3  |  Загрузок: 218

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.