K2-Horizon-MoVA-36B-A4B-APEX-GGUF
Imatrix-guided APEX quantization of IFM/K2-Horizon-MoVA-36B-A4B — MBZUAI’s Institute of Foundation Models (the LLM360/K2 lineage), released 2026-09-01, Apache-2.0. Not...
Imatrix-guided APEX quantization of IFM/K2-Horizon-MoVA-36B-A4B — MBZUAI’s Institute of Foundation Models (the LLM360/K2 lineage), released 2026-09-01, Apache-2.0. Not...
EXL3 квантование orcarouter/Qwen3.8-27B-Uncensored, аблитерированная (отказ-удаленная) сборка Qwen/Qwen3.8-27B. 16 ГБ на диске — подходит для одной карты 24 ГБ...
This experimental text-only build compresses Qwen3.8-Flash-Next from a 72.55 GB calibrated source GGUF to 39,721,239,200 bytes: — 36.993287...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
> Built with mlx-optiq, the MLX-native toolkit to quantize, prune, fine-tune, and serve LLMs locally on Apple Silicon....
Spark is a quality-first 2.80 BPW GGUF quant of zai-org/GLM-5.3-Flash, tuned to fit and run fully on a...
GGUF quantizations of GestaltLabs/Ornstein3.8-27B: a Qwen 3.8 27B dense vision-language fine-tune with interleaved linear and full attention (Gated...
GGUF quantizations of zai-org/GLM-5.3-Flash, made with llama.cpp. 320B total parameters, 18B active. 45 layers with hybrid attention: 34...
Two quantizations of the DFlash 2 draft model for Qwen/Qwen3.8-27B, both calibrated with an importance matrix captured from...
A second quantization of inclusionAI/Ling-3.0-flash, alongside Ling-3.0-flash-HybridQuant-NVFP4-W4A16. Identical bit placement and identical serving contract; it differs in how...