Original model: https://huggingface.co/nothingiisreal/L3.1-8B-Celeste-V1.5 Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset Thank you ZeroWw for the inspiration to experiment with embed/output If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run: You can either specify a new local-dir (L3.1-8B-Celeste-V1.5-Q8_0) or download them all in place (./) A great write up with charts showing various performances is provided by Artefact2 here The first thing to figure out is how big a model you can run. To do this, you’ll need to figure out how much RAM and/or VRAM you have. If you want your model running as FAST as possible, you’ll want to fit the whole thing on your GPU’s VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU’s total VRAM. If you want the absolute maximum quality, add both your system RAM and your GPU’s VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total. Next, you’ll need to decide if you want to use an ‘I-quant’ or a ‘K-quant’. If you don’t want to think too much, grab one of the K-quants. These are in…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: bartowski
Теги: gguf, not-for-all-audiences, en, endpoints_compatible, conversational
Лайков: 3 | Загрузок: 2,319
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.