GLM-5.2-504B-W4A16
> A 34%-expert-pruned GLM-5.2 that holds parity with the full unpruned model on a > well-powered real-world eval...
> A 34%-expert-pruned GLM-5.2 that holds parity with the full unpruned model on a > well-powered real-world eval...
NVFP4 (NVIDIA FP4, weight-only NVFP4A16) quantization of empero-ai/Qwythos-9B-Claude-Mythos-5-1M — a Claude-Mythos/Fable-trace reasoning fine-tune of Qwen3.5-9B (qwen35: a dense...
NVFP4-quantized build of prefeitura-rio/Rio-3.5-Open-397B — a 397B-parameter (17B active) Qwen3.5-MoE vision-language model (512 experts, hybrid softmax + linear/DeltaNet...
Этот репозиторий содержит черновик модели Gemma 4 31B Instruction-Tuned Assistant, квантованный до собственной точности FP4 (NVFP4) для высокоэффективного...
Этот репозиторий содержит настраиваемую по инструкциям модель Gemma 4 31B, квантованную с нативной точностью FP4 (NVFP4) для высокоэффективного...
NVFP4 GGUF quantizations of google/gemma-4-12B-it, for use with llama.cpp. The dense FFN tensors (all 48 layers × 3...
NVFP4 (W4A4) quantization of huihui-ai/Huihui-LFM2.5-8B-A1B-abliterated — the abliterated (refusal-reduced) build of Liquid AI’s 8.3B-total / 1.5B-active mixture-of-experts reasoner...
Весовые коэффициенты идентичны байтам вышестоящего количества (проверено config.json и model.safetensors.index.json SHA-256). Это зеркало существует для обеспечения закрепленной, стабильной,...
I’ve been trying to get as much performance as possible out of a single DGX Spark with Gemma...
— Used mmangkad/Qwen3.6-27B-NVFP4 as base model (Thank you for doing ModelOpt quant!) — Another mixed precision quant: ssmout...