Original model: https://huggingface.co/ifable/gemma-2-Ifable-9B All quants were made using the imatrix option (except BF16, that’s the original precision). The imatrix was generated with the dataset from here, using the BF16 GGUF with a context size of 8192 tokens (default is 512 but higher/same as model context size should improve quality) and 13 chunks. https://github.com/ggerganov/llama.cpp/tree/master/examples/imatrix https://github.com/ggerganov/llama.cpp/tree/master/examples/quantize
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: Hampetiudo
Теги: gguf, endpoints_compatible, conversational
Лайков: 3 | Загрузок: 1,806
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.