RedHatAI/Qwen3-14B-FP8-dynamic - Каталог нейросетей
Генерация текста

RedHatAI/Qwen3-14B-FP8-dynamic

Добавлено:
RedHatAI/Qwen3-14B-FP8-dynamic

— Model Architecture: Qwen3ForCausalLM — Input: Text — Output: Text — Model Optimizations: — Activation quantization: FP8 — Weight quantization: FP8 — Intended Use Cases: — Reasoning. — Function calling. — Subject matter experts via fine-tuning. — Multilingual instruction following. — Translation. — Out-of-scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). — Release Date: 05/02/2025 — Version: 1.0 — Model Developers: RedHat (Neural Magic) This model was obtained by quantizing activations and weights of Qwen3-14B to FP8 data type. This optimization reduces the number of bits used to represent weights and activations from 16 to 8, reducing GPU memory requirements (by approximately 50%) and increasing matrix-multiply compute throughput (by approximately 2x). Weight quantization also reduces disk size requirements by approximately 50%. Only weights and activations of the linear operators within transformers blocks are quantized. Weights are quantized with a symmetric static per-channel scheme, whereas activations are quantized with a symmetric dynamic per-token scheme. The llm-compressor library is used for quantization. This model…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: RedHatAI
Теги: qwen3, neuralmagic, redhat, llmcompressor, quantized, FP8, conversational, text-generation-inference
Лайков: 4  |  Загрузок: 1,113

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.