rushyrush/MiniMax-M2.1-REAP-172B-A10B-MXFP4_MOE-GGUF - Каталог нейросетей
Генерация текста

rushyrush/MiniMax-M2.1-REAP-172B-A10B-MXFP4_MOE-GGUF

Добавлено:
rushyrush/MiniMax-M2.1-REAP-172B-A10B-MXFP4_MOE-GGUF

This is a GGUF conversion and MXFP4 quantization of the model cerebras/MiniMax-M2.1-REAP-172B-A10B Original model: https://huggingface.co/cerebras/MiniMax-M2.1-REAP-172B-A10B model description from the creator: 𓌳 REAP𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression Introducing MiniMax-M2.1-REAP-172B-A10B, a memory-efficient compressed variant of MiniMax-M2.1 that maintains near-identical performance while being 25% lighter. This model was created using REAP (Router-weighted Expert Activation Pruning), a novel expert pruning method that selectively removes redundant experts while preserving the router’s independent control over remaining experts. Key features include: — Near-Lossless Performance: Maintains almost identical accuracy on code generation, agentic coding, and function calling tasks compared to the full 230B model — 25% Memory Reduction: Compressed from 230B to 172B parameters, significantly lowering deployment costs and memory requirements — Preserved Capabilities: Retains all core functionalities including code generation, math & reasoning and tool calling. — Drop-in Compatibility: Works with vanilla vLLM — no source modifications or custom patches…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: rushyrush
Теги: gguf, minimax, MOE, pruning, compression, en, endpoints_compatible, conversational
Лайков: 4  |  Загрузок: 13

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.