> Same size as a standard Q4KM. Measurably closer to the Q8 teacher across every measured lane (code/math, chat, tool calling, long-form text). !Fidelity receipts chart A drop-in Q4KM for Qwen 3.6 35B-A3B. Identical file size, identical kernel path, identical loader. ~23% lower output-distribution divergence from the Q8 teacher on code/math, ~8% lower on general (chat + tool calling + long-form text) vs a public Q4KM baseline — and ~42% / ~29% lower vs a public IQ4_XS baseline — measured on the same held-out slices, same prompts, same temperature. No retraining. No custom runtime. Standard llama.cpp Q4KM kernel. The win is in the calibration and per-tensor bit allocation. > 💻 This is the local / consumer ship. Runs on Apple Silicon (M-series) and consumer GPUs with stock llama.cpp — no patched runtime, no special flags. A separate MTP runtime variant targets datacenter speculative decoding (its 1.49× decode speedup is A100-80GB only and shows no speedup on consumer hardware), so for local use, this is the build you want. Same Q4KM file size, same llama.cpp kernel path, measured against two leading public Q4-class quants of the same base model. All five metrics run on the same eval…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: fraQtl
Теги: gguf, quantized, q4_k_m, qwen3.6, qwen3.6-35b-a3b, moe, llama-cpp, fraqtl
Лайков: 4 | Загрузок: 1,861
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.