This is a GGUF conversion and MXFP4 quantization of the model cerebras/MiniMax-M2.1-REAP-139B-A10B Original model: https://huggingface.co/cerebras/MiniMax-M2.1-REAP-139B-A10B 𓌳 REAP𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression Introducing MiniMax-M2.1-REAP-139B-A10B, a memory-efficient compressed variant of MiniMax-M2.1 that maintains near-identical performance while being 40% lighter. This model was created using REAP (Router-weighted Expert Activation Pruning), a novel expert pruning method that selectively removes redundant experts while preserving the router’s independent control over remaining experts. Key features include: — Near-Lossless Performance: Maintains almost identical accuracy on code generation, agentic coding, and function calling tasks compared to the full 230B model — 40% Memory Reduction: Compressed from 230B to 139B parameters, significantly lowering deployment costs and memory requirements — Preserved Capabilities: Retains all core functionalities including code generation, math & reasoning and tool calling. — Drop-in Compatibility: Works with vanilla vLLM — no source modifications or custom patches required — Optimized for Real-World…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: rushyrush
Теги: gguf, minimax, MOE, pruning, compression, en, endpoints_compatible, conversational
Лайков: 4 | Загрузок: 21
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.