— NOT Updated for new pre-tokenizer fixes (yet), I recommend using bartowski’s quants. https://huggingface.co/bartowski/Meta-Llama-3-70B-Instruct-GGUF — quants done with an importance matrix for improved quantization loss — K & IQ quants in basically all variants — fixed end token for instruct mode ([128009]) — files larger than 50GB were split using the gguf-split utility, just download all parts and point llama.cpp to the first one (00001-of-x) Quantized with llama.cpp commit with tokenizer fixes from this branch cherry-picked 0d56246f4b9764158525d894b96606f6 Therefore I have manually set the eos token to 128009 for these quants. In my testing this works fine, provide you you make sure to use the correct chat template. I recommend launching llama.cpp with —chat-template llama3 (make sure to use a newish version which has the PR for this merged). Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text…
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: qwp4w3hyb
Теги: gguf, facebook, meta, llama, llama-3, imatrix, importance matrix, conversational
Лайков: 3 | Загрузок: 936
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.