> 👋 I’m back. Active development has resumed — expect improvements soon. > ⚠️ Experimental. ASHQ1 is a personal research project that I will be refining over time. Use at your own risk. Results may vary between architectures and fine-tunes. Feedback and contributions welcome. Latest update (v7): MTP detection now catches the whole head (not just nextn.), output/embd are pinned at Q5K outside the budget, —allow-q3-or-lower actually covers its documented types, sub-4-bit toxicity (TOXICITYSUB4) fights the low-bit disease, MoE routers are pinned at F16, and there’s a new —top-down mode that starts everything at F16 and downgrades. New architectures: granite (Granite-4.2) and bailingmoe3 (Ling-3.0 MoE), detected via general.architecture metadata. Tested on Ornith-1.5-9B: PPL 8.6341 @ 6511 MiB. ASHQ1 is a post-training quantization method for GGUF models that uses an imatrix-driven priority queue to maximise theoretical quality per megabyte. Instead of uniform bit-depth or heuristic layer-blocking, it treats tied tensor groups as monolithic entities and greedily upgrades them by strict mathematical utility — the product of summed importance and theoretical MSE reduction, divided by…
Модальности:
Генерация текста
Задача: Генерация текста
Автор: wepiqx
Теги: gguf, quantization, llama-cpp, imatrix, hybrid-quantization, selective-quantization, priority-queue, mse
Лайков: 4 | Загрузок: 0
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.