Here are Quantized versions of Llama-3.1-70B-Instruct using GGUF GGUF is designed for use with GGML and other executors. GGUF was developed by @ggerganov who is also the developer of llama.cpp, a popular C/C++ LLM inference framework. Models initially developed in frameworks like PyTorch can be converted to GGUF format for use with those engines. — [ ] Q2K — [x] Q3KS — [x] Q3KM — [x] Q3KL — [x] Q4KS — [x] Q4KM ~ Recommended — [x] Q5KS ~ Recommended — [x] Q5KM ~ Recommended — [ ] Q6K — [ ] Q80 ~ NOT Recommended — [ ] F16 ~ NOT Recommended — [ ] F32 ~ NOT Recommended* Feel Free to reach out to me if you need a specific Quantization Type that I do not currently offer. Below is a table of all the Quantization Types that are possible as well as short descriptions. By using a GGUF version of Llama-3.1-70B-Instruct, you will be able to run this LLM while having to use significantly less resources than you would using the non quantized version. This also allows you to run this 70B Model on a machine with less memory than a non quantized version. Here are 2 different methods you can use to run the quantized versions of Llama-3.1-70B-Instruct Text-generation-webui is a web UI for Large…
Модальности:
Генерация текста
Области применения:
Следование инструкциям Диалог / чат
Задача: Генерация текста
Автор: hierholzer
Теги: gguf, llama, meta, llama-3.1, llama-3.1-instruct, ollama, Text-generation-webui, instruct
Лайков: 3 | Загрузок: 242
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.