mlx-community/DeepSeek-R1-Distill-Qwen-32B-MLX - Каталог нейросетей
Генерация текста

mlx-community/DeepSeek-R1-Distill-Qwen-32B-MLX

Добавлено:
mlx-community/DeepSeek-R1-Distill-Qwen-32B-MLX

This Model mlx-community/DeepSeek-R1-Distill-Qwen-32B contains multiple quantized variants of the base model deepseek-ai/DeepSeek-R1-Distill-Qwen-32B. The model was converted to MLX format using mlx-lm version 0.21.5. The conversion process applied different quantization strategies to produce variants that offer trade-offs between memory footprint, inference speed, and accuracy. In addition to the default 4-bit conversion, you will find both uniform and mixed quantized files at various bit widths (2-bit, 3-bit, 6-bit, and 8-bit). This multi-quantized approach allows users to select the best variant for their deployment scenario, balancing precision and performance. The model conversion uses a range of quantization configurations defined via mlxlm.convert`. These configurations fall into three main categories: 1. Uniform Quantization: Applies the same bit width to all layers. — 3bit: Uniform 3-bit quantization. — 4bit: Uniform 4-bit quantization (default). — 6bit: Uniform 6-bit quantization. — 8bit: Uniform 8-bit quantization. 2. Mixed Quantization: Uses a custom predicate function to decide the bit width for each layer—allowing different layers to use different precisions. -…

Модальности:
Генерация текста

Области применения:
Диалог / чат Логика и рассуждение


Задача: Генерация текста
Автор: mlx-community
Теги: mlx, chat, conversations, en
Лайков: 3  |  Загрузок: 0

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.