anikifoss/Kimi-K2-Instruct-DQ4_K - Каталог нейросетей
Генерация текста

anikifoss/Kimi-K2-Instruct-DQ4_K

Добавлено:
anikifoss/Kimi-K2-Instruct-DQ4_K

High quality quantization of Kimi-K2-Instruct without using imatrix. You may be able to run with 512G of RAM by removing —no-mmap and -rtr with a performance hit that will depend on what MoE experts get activated for your prompt. See this detailed guide on how to setup ik_llama and how to make custom quants. — Keep all the small F32 tensors untouched — Quantize all the attention and related tensors to Q80 — Quantize all the ffndownexps tensors to Q6K — Quantize all the ffnupexps and ffngateexps tensors to Q4K` Generally, imatrix is not recommended for Q4 and larger quants. The problem with imatrix is that it will guide what model remembers, while anything not covered by the text sample used to generate the imartrix is more likely to be forgotten. For example, an imatrix derived from wikipedia sample is likely to negatively affect tasks like coding. In other words, while imatrix can improve specific benchmarks, that are similar to the imatrix input sample, it will also skew the model performance towards tasks similar to the imatrix sample at the expense of other tasks.

Модальности:
Генерация текста

Области применения:
Диалог / чат Следование инструкциям


Задача: Генерация текста
Автор: anikifoss
Теги: gguf, mla, conversational, no_imatrix, endpoints_compatible
Лайков: 4  |  Загрузок: 65

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.