legraphista/RoLlama3-8b-Instruct-IMat-GGUF - Каталог нейросетей
Генерация текста

legraphista/RoLlama3-8b-Instruct-IMat-GGUF

Добавлено:
legraphista/RoLlama3-8b-Instruct-IMat-GGUF

Llama.cpp imatrix quantization of OpenLLM-Ro/RoLlama3-8b-Instruct Original Model: OpenLLM-Ro/RoLlama3-8b-Instruct Original dtype: BF16 (bfloat16) Quantized by: llama.cpp b3206 IMatrix dataset: here — Files — IMatrix — Common Quants — All Quants — Downloading using huggingface-cli — Inference — Simple chat template — Chat template with system prompt — Llama.cpp — FAQ — Why is the IMatrix not applied everywhere? — How do I merge a split GGUF? —— | ———- | ——— | —— | ———— | ——— | If the model file is big, it has been split into multiple files. In order to download them all to a local folder, run: According to this investigation, it appears that lower quantizations are the only ones that benefit from the imatrix input (as per hellaswag results). 1. Make sure you have gguf-split available — To get hold of gguf-split, navigate to https://github.com/ggerganov/llama.cpp/releases — Download the appropriate zip for your system from the latest release — Unzip the archive and you should be able to find gguf-split 2. Locate your GGUF chunks folder (ex: RoLlama3-8b-Instruct.Q80) 3. Run gguf-split —merge…

Модальности:
Генерация текста

Области применения:
Диалог / чат Следование инструкциям


Задача: Генерация текста
Автор: legraphista
Теги: gguf, quantized, GGUF, quantization, imat, imatrix, static, 16bit
Лайков: 3  |  Загрузок: 2,315

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.