The model is the GGUF version of rinna/nekomata-14b. It can be used with llama.cpp for lightweight inference. Quantization of this model may cause stability issue in GPTQ, AWQ and GGUF q40. We recommend GGUF q4KM** for 4-bit quantization. See rinna/nekomata-14b for details about model architecture and data. Please refer to rinna/nekomata-14b for tokenization details.
Модальности:
Генерация текста
Задача: Генерация текста
Автор: rinna
Теги: gguf, qwen, ja, en
Лайков: 3 | Загрузок: 48
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.