Sparse-Llama-3.1-8B-2of4-GGUF
This is quantized version of neuralmagic/Sparse-Llama-3.1-8B-2of4 created using llama.cpp — Model Architecture: Llama-3.1-8B — Input: Text — Output:...
This is quantized version of neuralmagic/Sparse-Llama-3.1-8B-2of4 created using llama.cpp — Model Architecture: Llama-3.1-8B — Input: Text — Output:...
В моделях Sparse-Llama-3.1 используется полуструктурированная разреженность 2:4, что позволяет увеличить размер модели в 2 раза и сократить объем...