mlx-community/Ornith-1.5-35B-A3B-OptiQ-4bit-REAP-19B - Каталог нейросетей
Генерация текста

mlx-community/Ornith-1.5-35B-A3B-OptiQ-4bit-REAP-19B

Добавлено:
mlx-community/Ornith-1.5-35B-A3B-OptiQ-4bit-REAP-19B

> Built with mlx-optiq, the MLX-native toolkit to quantize, prune, fine-tune, and serve LLMs locally on Apple Silicon. All OptiQ models · Docs 50% of the routed experts are removed from mlx-community/Ornith-1.5-35B-A3B-OptiQ-4bit; active parameters per token are unchanged, since top-8 routing is preserved and only the stored expert bank shrinks. That is why it gets smaller without getting slower. Retained experts are copied bit-for-bit from the parent quant. Nothing is dequantized, re-quantized, merged, or retrained. This variant was not separately benchmarked. It is published under the recipe validated end to end on Qwen3.6-35B-A3B-OptiQ-4bit-REAP-19B, the same architecture at the same 50 % retention: Capability Score 80.03 -> 76.57, with the loss concentrated in MMLU (-21.4) and procedural ability intact (GSM8K +2.6, IFEval +4.3, BFCL -1.0, HumanEval -1.3). Two things were measured on this checkpoint. The ranking rule was chosen by scoring both candidates against the unpruned model, which picked the conditional mean. And the resulting divergence from the unpruned parent is KL 0.331 — for reference, the checkpoints that degrade visibly under pruning measure above 1.0, and this…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: mlx-community
Теги: mlx, qwen3_5_moe, quantized, expert-pruning, reap, moe, optiq, apple-silicon
Лайков: 4  |  Загрузок: 745

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.