Pantheon-RP-1.5-12b-Nemo-GGUF
Original model: https://huggingface.co/Gryphe/Pantheon-RP-1.5-12b-Nemo Some of these quants (Q3KXL, Q4KL etc) are the standard quantization method with the embeddings...
Original model: https://huggingface.co/Gryphe/Pantheon-RP-1.5-12b-Nemo Some of these quants (Q3KXL, Q4KL etc) are the standard quantization method with the embeddings...
My own (ZeroWw) quantizations. output and embed tensors quantized to f16. all other tensors quantized to q5k or...
Original model: https://huggingface.co/Nitral-AI/Hathor_Sofit-L3-8B-v1 Some of these quants (Q3KXL, Q4KL etc) are the standard quantization method with the embeddings...
Here are Quantized versions of Llama-3.1-70B-Instruct using GGUF GGUF is designed for use with GGML and other executors....
Original model: https://huggingface.co/NeverSleep/Lumimaid-v0.2-12B Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset Thank you...
Original model: https://huggingface.co/aifeifei798/DarkIdol-Llama-3.1-8B-Instruct-1.0-Uncensored Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset Thank you...
Original model: https://huggingface.co/nothingiisreal/L3.1-8B-Celeste-V1.5 Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset Thank you...
Using llama.cpp commit b5e9546 for quantization, featuring llama 3.1 rope scaling factors. This fixes low-quality issues when using...
ZeroWw ‘SILLY’ version. The original model has been quantized (fq8 version) and a percentage of it’s tensors have...
 Это квантованная версия jpacifico/Chocolatine-3B-Instruct-DPO-Revised, созданная с использованием llama.cpp DPO, точно настроенная на microsoft/Phi-3-mini-4k-instruct (3.82B params) с использованием...