!Format !Task !Params !Type !Quant !Size !Context !Refusals !KL drift !License > Yes — MTP PRESERVED fp16 inside this model — native multi-token-prediction speculative decoding works with no external drafter. Yes — VISION tower preserved fp16. Yes — SSM-sensitive params (alog, dtbias, conv1d) kept fp16. Quantized with mlx-mtp — a pure–Apple-mlx stack (no third-party ML-inference frameworks at runtime). MXFP8 (8-bit microscaling) MLX quantization of a abliterated Qwen 3.6 27B Coder (Jackrong’s agentic-coding SFT of Qwen 3.6 27B × Claude-Opus reasoning distill). Refusals reduced from 86/100 → 8/100 with KL drift of 0.007. Tensor set is identical to the base model (1199 tensors: 333 vision + 15 MTP). By the Lemura Labs research team. Measured with the ablation toolkit on mlabonne/harmfulbehaviors (100 hard red-team prompts) and KL divergence on mlabonne/harmlessalpaca. → 90.7% reduction in refusals with coding capabilities preserved. No SFT / LoRA healing required. mlx-mtp loads this checkpoint natively (vision + MTP), on Apple mlx only: We are deeply grateful to everyone whose work made this release possible. Foundation Model — Qwen Team @ Alibaba Tongyi Lab, for Qwen3.6-27B: a…
Модальности:
Генерация текста Компьютерное зрение Мультимодальность
Области применения:
Логика и рассуждение Диалог / чат Вызов функций (Tool use) Мультиязычность Генерация кода
Задача: Генерация текста
Автор: lemuralabs
Теги: mlx, qwen3_5, mlx-mtp, qwen, qwen3, qwen3.5, qwen3.6, claude-opus-distill
Лайков: 4 | Загрузок: 1,728
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.