GLM-5.2-Abliterated-MXFP8-NVFP4-NF3-Hybrid
This is a full-expert GLM-5.2 hybrid checkpoint that combines the compact MXFP8/NVFP4/NF3 layout from madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid with selected refusal-reduction...
This is a full-expert GLM-5.2 hybrid checkpoint that combines the compact MXFP8/NVFP4/NF3 layout from madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid with selected refusal-reduction...
MLX 4-bit affine quantization of poolside/Laguna-S-2.1, packaged for Apple Silicon experiments and local OpenAI-compatible serving. Laguna S 2.1...
A 33%-expert-pruned version of Tencent Hunyuan Hy3 (295B / A21B), compressed with REAP (Router-weighted Expert Activation Pruning) —...
Non-uniform expert prune · NVFP4 · runs on FOUR Blackwell cards · verified 0-refusal Three transformations in one...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
weighted/imatrix quants of https://huggingface.co/hotdogs/Qwen35B-Agent-R2 For a convenient overview and download list, visit our model page for this model....
Quantized tencent/Hy3 for Apple Silicon MLX / JANG runtimes — a 295B-total / 21B-active text MoE, packed to...
High-quality imatrix GGUF quantizations of tencent/Hy3, Tencent’s 295B-parameter Mixture-of-Experts model with ~21B active parameters per token. Produced with...
QLoRA fine-tune of NVIDIA Nemotron-3-Super-120B-A12B (120B-total / 12B-active hybrid Mamba-2 + Latent-MoE, nemotronh) on the piagent split of...
QLoRA fine-tune of NVIDIA Nemotron-3-Super-120B-A12B (120B-total / 12B-active hybrid Mamba-2 + Latent-MoE, nemotronh) on the piagent split of...