Ornith-1.5-35B-A3B-OptiQ-4bit-REAP-19B
> Built with mlx-optiq, the MLX-native toolkit to quantize, prune, fine-tune, and serve LLMs locally on Apple Silicon....
> Built with mlx-optiq, the MLX-native toolkit to quantize, prune, fine-tune, and serve LLMs locally on Apple Silicon....
A surgically expert-pruned openai/gpt-oss-120b: 48 of 128 experts kept per layer (top-48 by REAP saliency), reducing the checkpoint...
A 33%-expert-pruned version of Tencent Hunyuan Hy3 (295B / A21B), compressed with REAP (Router-weighted Expert Activation Pruning) —...
This repository contains custom APEX (Adaptive Precision for EXpert Models) -inspired quants, tuned specifically to the underlying mixture-of-experts...
> A 34%-expert-pruned GLM-5.2 that holds parity with the full unpruned model on a > well-powered real-world eval...
> [!TIP] > Support this work → · X · GitHub · REAP paper · Cerebras REAP 30%...
50% Экспертная обрезка — 80B → 24 ГБ при сохранении 93,5% оригинального качества. Половина всех экспертов MoE удалена...
This model was converted to GGUF format from AmanPriyanshu/gpt-oss-4.2b-specialized-all-pruned-moe-only-4-experts using llama.cpp via the ggml.ai’s GGUF-my-repo space. Refer to...
Project: https://amanpriyanshu.github.io/GPT-OSS-MoE-ExpertFingerprinting/ This is a pruned variant of OpenAI’s GPT-OSS-20B model, reduced to 4 experts per layer based...
Уменьшение числа экспертов-специалистов по коду из Qwen3.6-35B-A3B: количество экспертов MoE уменьшено с 256 до 184 (72 удалено на...