Qwen3.8-Flash-Next-REAM-60Pct-GGUF
GGUF quantizations of Akicou/Qwen3.8-Flash-Next-REAM-60Pct, the REAM-compressed (Merged) version of Qwen/Qwen3.8-Flash-Next. REAM (Router Expert Activation Merging) pruned 40% of...
GGUF quantizations of Akicou/Qwen3.8-Flash-Next-REAM-60Pct, the REAM-compressed (Merged) version of Qwen/Qwen3.8-Flash-Next. REAM (Router Expert Activation Merging) pruned 40% of...
A 33%-expert-pruned version of Tencent Hunyuan Hy3 (295B / A21B), compressed with REAP (Router-weighted Expert Activation Pruning) —...
This repository contains a REAM-compressed version of Tencent Hy3, produced with Akicou/ream — a REAM/REAP-style Mixture-of-Experts compression framework....
This is a GGUF conversion and MXFP4 quantization of the model cerebras/MiniMax-M2.1-REAP-172B-A10B Original model: https://huggingface.co/cerebras/MiniMax-M2.1-REAP-172B-A10B model description from...
This is a GGUF conversion and MXFP4 quantization of the model cerebras/MiniMax-M2.1-REAP-139B-A10B Original model: https://huggingface.co/cerebras/MiniMax-M2.1-REAP-139B-A10B 𓌳 REAP𓌳 the...
𓌳 REAP𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression Introducing GLM-4.7-Flash-REAP-23B-A3B, a memory-efficient compressed variant of...
𓌳 REAP𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression Introducing MiniMax-M2-REAP-162B-A10B, a memory-efficient compressed variant of...
𓌳 REAP𓌳 the Experts: Why Pruning Prevails for One-Shot MoE Compression Introducing GLM-4.6-REAP-252B-A32B, a memory-efficient compressed variant of...
𓌳 REAP𓌳 Эксперты: почему обрезка преобладает при однократном сжатии MoE Представляем Step-3.5-Flash-REAP-149B-A11B, сжатый вариант Step-3.5-Flash с эффективным использованием...
Эта модель представляет собой сжатую версию Qwen/Qwen3-Next-80B-A3B-Instruct. Его получают за счет сокращения количества экспертов на каждом уровне MoE...