FP8-quantized version of DeepSeek-R1-Distill-Qwen-14B, optimized for inference with vLLM. The quantization reduces the model’s memory footprint by approximately 50%. — Base Model: DeepSeek-R1-Distill-Qwen-14B — Quantization: FP8 (weights and activations) — Memory Reduction: ~50% (from 16-bit to 8-bit) — License: MIT License (following original model’s license) — 512 calibration samples from UltraChat — Symmetric per-tensor quantization — Applied to linear operators within transformer blocks This is an experimental compression of the model. Performance metrics and optimal usage parameters have not been thoroughly tested yet.
Модальности:
Генерация текста
Области применения:
Диалог / чат Логика и рассуждение
Задача: Генерация текста
Автор: enferAI
Теги: qwen2, fp8, vllm, conversational, en, de, fr, it
Лайков: 3 | Загрузок: 27
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.