— Model Architecture: Qwen2ForCausalLM — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 — Activation quantization: FP8 — Release Date: 2/5/2025 — Version: 1.0 — Model Developers: Neural Magic This model was obtained by quantizing the weights and activations of DeepSeek-R1-Distill-Qwen-14B to FP8 data type. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. Only the weights and activations of the linear operators within transformers blocks are quantized. Weights are quantized using a symmetric per-channel scheme, whereas quantizations are quantized using a symmetric per-token scheme. LLM Compressor is used for quantization. This model can be deployed efficiently using the vLLM backend, as shown in the example below. vLLM also supports OpenAI-compatible serving. See the documentation for more details. This model was created with llm-compressor by running the code snippet below. The model was evaluated on OpenLLM Leaderboard V1 and V2, using the following commands: Category Metric deepseek-ai/DeepSeek-R1-Distill-Qwen-14B…
Модальности:
Генерация текста
Области применения:
Диалог / чат Логика и рассуждение
Задача: Генерация текста
Автор: RedHatAI
Теги: qwen2, deepseek, fp8, vllm, conversational, text-generation-inference, endpoints_compatible, compressed-tensors
Лайков: 3 | Загрузок: 525
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.