Original model: https://huggingface.co/ValiantLabs/Llama3.1-8B-Enigma Some of these quants (Q3KXL, Q4KL etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to. Some say that this improves the quality, others don’t notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don’t keep uploading quants no one is using. Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset Thank you ZeroWw for the inspiration to experiment with embed/output If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run: You can either specify a new local-dir (Llama3.1-8B-Enigma-Q8_0) or download them all in place (./) A great write up with charts showing various performances is provided by Artefact2 here The first thing to figure out is how big a model you can run. To do this, you’ll need to figure out how much RAM and/or VRAM you have. If you want your model running as FAST as possible, you’ll want to fit the whole thing…
Модальности:
Генерация текста
Области применения:
Генерация кода Следование инструкциям Диалог / чат
Языки программирования:
Python
Задача: Генерация текста
Автор: bartowski
Теги: gguf, enigma, valiant, valiant-labs, llama, llama-3.1, llama-3.1-instruct, llama-3.1-instruct-8b
Лайков: 3 | Загрузок: 2,091
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.