GLM-5.3-Flash-Spark-Q2XL-MTP
Spark is a quality-first 2.80 BPW GGUF quant of zai-org/GLM-5.3-Flash, tuned to fit and run fully on a...
Spark is a quality-first 2.80 BPW GGUF quant of zai-org/GLM-5.3-Flash, tuned to fit and run fully on a...
GGUF quantizations of GestaltLabs/Ornstein3.8-27B: a Qwen 3.8 27B dense vision-language fine-tune with interleaved linear and full attention (Gated...
GGUF quantizations of zai-org/GLM-5.3-Flash, made with llama.cpp. 320B total parameters, 18B active. 45 layers with hybrid attention: 34...
Two quantizations of the DFlash 2 draft model for Qwen/Qwen3.8-27B, both calibrated with an importance matrix captured from...
A second quantization of inclusionAI/Ling-3.0-flash, alongside Ling-3.0-flash-HybridQuant-NVFP4-W4A16. Identical bit placement and identical serving contract; it differs in how...
An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language...
4-bit GPTQ (W4A16) quantization of deepseek-ai/DeepSeek-V4-Flash-0731, produced with GPTQModel. This is the V1 release (first calibration run). A...
> Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no...
Unofficial GGUF conversion and importance-matrix quantizations of inclusionAI/Ling-3.0-tiny, created from immutable source revision a2ee06c0. No fine-tuning, merging, or...
GGUF квантовала версии inclusionAI/Ling-3.0-flash, нативной гибридной модели рассуждений следующего поколения с общими параметрами 124B и активными параметрами 5.1B...