I have uploaded what I believe to be the final versions of the quantized models, created with a strict maximum size target of 14.1 GB (without MTP). Qwen3.6-27B.i1-IQ4KT-attnqkv-IQ4KS.gguf (no MTP, for GPU Only) Qwen3.6-27B.i1-IQ4KT-attnqkv-IQ4KS-i1MTP.gguf (with MTP, for GPU Only) Qwen3.6-27B.CPU.i1-IQ4KT-attnqkv-IQ4KS-i1MTP.gguf (with MTP, for GPU+CPU) The most interesting build here is Qwen3.6-27B.CPU.i1-IQ4KT-attnqkv-IQ4KS-i1MTP.gguf, which uses iq4ks` quantization for the first 16 FFN blocks. This was done to test out the concept discussed in this Reddit thread. GPU/CPU setup will give you 140k ctx (q50/q40) with decode speed starting at 35-30t/s and fall to 15-20t/s at the end (RTX 5070Ti + 8845HS dual channel DDR 5600) IMPORTAN: Somehow the ubatch-size is important for the quality of this model. The recommended value is —ubatch-size 192. INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED BUT SOME MAY HAVE NO CONTENT
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: cHunter789
Теги: gguf, qwen, ik_llama.cpp, nvidia, imatrix, 4-bit, iq4_ks, endpoints_compatible
Лайков: 4 | Загрузок: 2,371
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.