These are iqk format quantizations and are EXCLUSIVE to the ikllama.cpp fork.** They will NOT work on mainline llama.cpp, standard LM Studio, standard Text Generation WebUI, or KoboldCPP. You must compile and run this using ikawrakow’s llama.cpp fork, or a UI where you have manually swapped the backend to an ikllama.cpp` build. All variants use the same custom tensor buckets: attention, SSM, shared experts, routed experts, embeddings/output, and MTP/NextN tensors. High quality routed expert quantization using IQ5KS` experts. Balanced 4-bit routed expert quantization with high precision on always-active tensors. Smaller 4-bit routed expert quantization with compressed embeddings, output, and MTP tensors. Lower size recipe with IQ3K routed experts and IQ6K on many always-active tensors. 1. Clone and build the ikllama.cpp fork from ikawrakow/ikllama.cpp. 2. Use the compiled llama-server or llama-cli from that specific build. 3. For chat templating, use the model’s embedded template or the community template credited above, depending on your frontend.
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: KeinNiemand
Теги: gguf, quantization, iq6_k, iq5_k, iq5_ks, iq4_k, iq4_ks, iq4_kss
Лайков: 4 | Загрузок: 198
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.