Qwen3.6-35B-A3B-MTP-4bit
This repository contains quantized Multi-Token Prediction (MTP) drafter weights split from Qwen/Qwen3.6-35B-A3B for use with mlx-vlm speculative decoding....
This repository contains quantized Multi-Token Prediction (MTP) drafter weights split from Qwen/Qwen3.6-35B-A3B for use with mlx-vlm speculative decoding....
MagicQuant is a benchmark driven GGUF hybrid discovery and validation system focused on finding real, practical GGUF quants...
35B MoE Qwen3.6-A3B trunk (3B active) + embedded NextN-MTP head, quantized for single-GPU inference. — Trunk: IQ4XS (imatrix-calibrated)...
MagicQuant is a benchmark driven GGUF hybrid discovery and validation system focused on finding real, practical GGUF quants...
Hermes-style agentic fine-tune of Qwen3.6-27B, quantized to INT4 with a BF16 MTP overlay for speculative decoding. This model...
MLX conversion of Qwen/Qwen3.6-27B, affine 8-bit (group_size 64), with the native Multi-Token-Prediction (MTP) head embedded in the main...
First public extraction of Google Gemma 4’s Multi-Token Prediction (MTP) drafter weights from LiteRT format into standard PyTorch...
Chadrockv2 Qwen3.6 27B ROCmFP6 STRIX QUALITY — это настроенная AMD версия GGUF линейки Unsloth Qwen3.6 27B MTP. Он...
> MTP дает ~2 × одиночный поток (базовая скорость 7,7 → 15 ток/с, до ~19 / акцепт 0,89...
Отдельный проект MTP (многотокеновое предсказание) предназначен для спекулятивного декодирования с помощью CosmicRaisins/GLM-5.2-AWQ-INT4-15pct. cyankiwi/GLM-5.2-AWQ-INT4 удаляет собственный уровень MTP GLM-5.2,...