Llama3-ChatQA-2-8B-GGUF
This is quantized version of nvidia/Llama3-ChatQA-2-8B created using llama.cpp We introduce Llama3-ChatQA-2, a suite of 128K long-context models,...
This is quantized version of nvidia/Llama3-ChatQA-2-8B created using llama.cpp We introduce Llama3-ChatQA-2, a suite of 128K long-context models,...
This is quantized version of rinna/gemma-2-baku-2b-it created using llama.cpp The model is an instruction-tuned variant of rinna/gemma-2-baku-2b, utilizing...
This model was converted to GGUF format from KBTG-Labs/THaLLE-0.1-7B-fa using llama.cpp. Refer to the original model card for...
This is quantized version of google/gemma-2-2b-jpn-it created using llama.cpp — Responsible Generative AI Toolkit — Gemma 2 JPN...
This is quantized version of universitytehran/PersianMind-v1.0 created using llama.cpp PersianMind is a cross-lingual Persian-English large language model. The...
This model was converted to GGUF format from openthaigpt/openthaigpt1.5-7b-instruct using llama.cpp via the ggml.ai’s GGUF-my-repo space. Refer to...
Original model: https://huggingface.co/EVA-UNIT-01/EVA-Qwen2.5-14B-v0.0 Some of these quants (Q3KXL, Q4KL etc) are the standard quantization method with the embeddings...
This model was converted to GGUF format from pcuenq/Qwen2.5-0.5B-Instruct-with-new-merges-serialization using llama.cpp via the ggml.ai’s GGUF-my-repo space. Refer to...
Original model: https://huggingface.co/TheDrummer/Cydonia-22B-v1.1 Some of these quants (Q3KXL, Q4KL etc) are the standard quantization method with the embeddings...
Some of these quants (Q3KXL, Q4KL etc) are the standard quantization method with the embeddings and output weights...