This works well as a draft model for speculative decoding in LMstudio 3.10 beta Try it with: mlx-community/FuseO1-DeepSeekR1-Qwen2.5-Coder-32B-4.5bit you should see 30% faster TPS for math/code prompts even with «thinking» slowing down the Specultive Decoding The Model bobig/DeepScaleR-1.5B-6.5bit was converted to MLX format from agentica-org/DeepScaleR-1.5B-Preview using mlx-lm version 0.21.4.
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: mlx-community
Теги: qwen2, mlx, conversational, en, text-generation-inference, endpoints_compatible, 6-bit
Лайков: 4 | Загрузок: 52
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.