This quant collection REQUIRES ikllama.cpp fork to support advanced non-linear SotA quants. Do not** download these big files and expect them to run on mainline vanilla llama.cpp, ollama, LM Studio, KoboldCpp, etc! NOTE ikllama.cpp` can also run your existing GGUFs from bartowski, unsloth, mradermacher, etc if you want to try it out before downloading my quants. These quants provide best in class quality for the given memory footprint. Shout out to Wendell and the Level1Techs crew, the community Forums, YouTube Channel! BIG thanks for providing BIG hardware expertise and access to run these experiments and make these great quants available to the community!!! Also thanks to all the folks in the quanting and inferencing community here and on r/LocalLLaMA for tips and tricks helping each other run all the fun new models! So far these are my best recipes offering the great quality in good memory footprint breakpoints. — type f32: 161 tensors — norms etc. — type iq6k: 2 tensors — tokenembd/output — type iq4ks: 80 tensors — ffn(gate|up) — type iq5ks: 200 tensors — ffndown and all attn This quant is designed to take advantage of faster iq4ks and new iq5ks quants. This quant…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: ubergarm
Теги: gguf, imatrix, conversational, ik_llama.cpp, endpoints_compatible
Лайков: 4 | Загрузок: 23
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.