GLM-5.3-MXFP4
— Model architecture: full GLM-5.3 (GlmMoeDsaForCausalLM) — Input: text — Output: text — Source checkpoint: zai-org/GLM-5.3-BF16, revision 304b8051cfb2b260b61ce0cbe330e02a98e73639...
— Model architecture: full GLM-5.3 (GlmMoeDsaForCausalLM) — Input: text — Output: text — Source checkpoint: zai-org/GLM-5.3-BF16, revision 304b8051cfb2b260b61ce0cbe330e02a98e73639...
This is an MXFP4 MLX quantization of JetBrains/Mellum2-12B-A2.5B-Instruct. Mellum2 Instruct is a Mixture-of-Experts assistant model with 64 experts...
A surgically expert-pruned openai/gpt-oss-120b: 48 of 128 experts kept per layer (top-48 by REAP saliency), reducing the checkpoint...
Специализированные нецензурированные/аблитерированные кванты для новой модели OpenAI 20B MOE — Смесь экспертов при 80 т/с. См. настройки и...
GGUF conversions of the tiny Kimi-K3 0.40B development checkpoints. — Kimi-K3 architecture validation; — llama.cpp conversion and inference...
TurboQuant + MLX-MXFP4 (4-bit) variant of Qwen/Qwen3.6-35B-A3B. Apple Silicon (M1/M2/M3/M4) with TurboQuant structural pre-conditioning and MLX-native MXFP4 layout...
A refusal-suppressed variant of openai/gpt-oss-20b, produced with abliterix using direct weight editing, Expert-Granular Abliteration (EGA) on the fused...
Когда вы хотите специфическое для MoE сжатие без отправки качества в бездну. Построено из: — Основание: MiniMaxAI/MiniMax-M2.5 —...
(DEPRECIATED — Part of MagicQuant v1.0 which had significant flaws. Please utilize v2.0 which is production ready) >...
This is an improved version using RL of the vibe-code LLM. It’s optimized to produce both natural-language and...