blockchainlabs_7B_merged_test2_4_prune
blockchainlabs7Bmergedtest24prune — это сокращенная модель, основанная на alnrg2arg/blockchainlabs7Bmergedtest24, которая представляет собой объединенную модель с использованием следующих моделей с...
blockchainlabs7Bmergedtest24prune — это сокращенная модель, основанная на alnrg2arg/blockchainlabs7Bmergedtest24, которая представляет собой объединенную модель с использованием следующих моделей с...
This repository contains custom APEX (Adaptive Precision for EXpert Models) -inspired quants, tuned specifically to the underlying mixture-of-experts...
> A 34%-expert-pruned GLM-5.2 that holds parity with the full unpruned model on a > well-powered real-world eval...
static quants of https://huggingface.co/dervig/m51Lab-MiniMax-M2.7-REAP-139B-A10B For a convenient overview and download list, visit our model page for this model....
50% Экспертная обрезка — 80B → 24 ГБ при сохранении 93,5% оригинального качества. Половина всех экспертов MoE удалена...
This is a GGUF conversion and MXFP4 quantization of the model cerebras/MiniMax-M2.1-REAP-172B-A10B Original model: https://huggingface.co/cerebras/MiniMax-M2.1-REAP-172B-A10B model description from...
This is a GGUF conversion and MXFP4 quantization of the model cerebras/MiniMax-M2.1-REAP-139B-A10B Original model: https://huggingface.co/cerebras/MiniMax-M2.1-REAP-139B-A10B 𓌳 REAP𓌳 the...
𓌳 REAP𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression Introducing GLM-4.7-Flash-REAP-23B-A3B, a memory-efficient compressed variant of...
𓌳 REAP𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression Introducing MiniMax-M2-REAP-162B-A10B, a memory-efficient compressed variant of...
𓌳 REAP𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression Introducing GLM-4.6-REAP-252B-A32B, a memory-efficient compressed variant of...