This repository contains quantized Multi-Token Prediction (MTP) drafter weights split from Qwen/Qwen3.6-35B-A3B for use with mlx-vlm speculative decoding. This is not a standalone chat or text-generation model. Load it as the draft model alongside a compatible Qwen3.6 35B-A3B target checkpoint. — Model type: qwen35mtp — MTP block size: 2 — Target architecture: Qwen3.6 35B-A3B — Precision: MLX affine 4-bit, group size 64 — Runtime: MLX / mlx-vlm — Format: Safetensors with MLX-compatible config and tokenizer files The stored tensors use MLX affine quantization as described in config.json. Use this repo only as a speculative decoding drafter for compatible Qwen3.6 35B-A3B checkpoints. The target model verifies drafted tokens, while this MTP model proposes candidate tokens per decoding step. This checkpoint requires runtime support for Qwen MTP draft models in mlx-vlm. Standard standalone generation through generic Transformers APIs is not expected to work with this repository by itself. Please refer to the upstream Qwen/Qwen3.6-35B-A3B model card and license terms for model usage constraints.
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: mlx-community
Теги: mlx, qwen3_5_mtp, mlx-vlm, qwen3.6, qwen3.6-35b-a3b, qwen, moe, mtp
Лайков: 4 | Загрузок: 1,828
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.