Akicou/Qwen3.8-Flash-Next-REAM-60Pct-GGUF - Каталог нейросетей
Генерация текста

Akicou/Qwen3.8-Flash-Next-REAM-60Pct-GGUF

Добавлено:
Akicou/Qwen3.8-Flash-Next-REAM-60Pct-GGUF

GGUF quantizations of Akicou/Qwen3.8-Flash-Next-REAM-60Pct, the REAM-compressed (Merged) version of Qwen/Qwen3.8-Flash-Next. REAM (Router Expert Activation Merging) pruned 40% of the routed experts in the original model, taking each layer from 512 down to 308 experts. The compressed checkpoint was then converted to GGUF with ggml-org/llama.cpp (converthftogguf.py, bf16) and quantized with llama-quantize`. No importance matrix was used. The architecture is qwen4exp (hybrid linear attention + Qwen Sparse Attention MoE), 48 layers, 308 routed experts per layer. — Experimental release, not benchmarked. — The base model requires trustremotecode=True. These GGUF files are for llama.cpp (and compatible runtimes), so remote code is not needed at load. — Shared experts, attention, and n-gram embeddings are untouched; only routed experts were merged.

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: Akicou
Теги: gguf, qwen, ream, merged, compression, mixture-of-experts, moe, llama.cpp
Лайков: 4  |  Загрузок: 514

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.