0xSero/MiniMax-M2.1-162B - Каталог нейросетей
Генерация текста

0xSero/MiniMax-M2.1-162B

Добавлено:
0xSero/MiniMax-M2.1-162B

> [!TIP] > Support this work → · X · GitHub · REAP paper · Cerebras REAP 30% expert-pruned MiniMax-M2.1 using REAP (Router-weighted Expert Activation Pruning) Tested at 4 temperatures (0.0, 0.2, 0.7, 1.0) across 6 prompt types (24 total tests): Additional tests at temperatures 0.5, 0.8, 0.9, 1.2 (results in stresstestresults.json). If you encounter TypeError: CacheLayerMixin.init() got an unexpected keyword argument, add this before importing the model: — MiniMax-M2.1-REAP-40-W4A16 (Coming Soon) — 4-bit weights, ~58GB REAP (Router-weighted Expert Activation Pruning) uses calibration data to identify which experts are most important based on router activation patterns. Unlike random or magnitude-based pruning, REAP preserves the experts that are actually used during inference. Calibration Dataset: 2098 samples — pile-10k: 498 samples (general text) — evol-codealpaca: 800 samples (code generation) — xlam-function-calling: 800 samples (function calling) Made possible by NVIDIA · TNG Technology · Lambda · Prime Intellect · Hot Aisle.

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: 0xSero
Теги: minimax_m2, minimax, moe, pruned, reap, conversational, custom_code, endpoints_compatible
Лайков: 4  |  Загрузок: 33

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.