Ornith-1.5-35B-A3B-OptiQ-4bit-REAP-19B
> Built with mlx-optiq, the MLX-native toolkit to quantize, prune, fine-tune, and serve LLMs locally on Apple Silicon....
> Built with mlx-optiq, the MLX-native toolkit to quantize, prune, fine-tune, and serve LLMs locally on Apple Silicon....
A surgically expert-pruned openai/gpt-oss-120b: 48 of 128 experts kept per layer (top-48 by REAP saliency), reducing the checkpoint...
A 33%-expert-pruned version of Tencent Hunyuan Hy3 (295B / A21B), compressed with REAP (Router-weighted Expert Activation Pruning) —...
This repository contains a REAM-compressed version of Tencent Hy3, produced with Akicou/ream — a REAM/REAP-style Mixture-of-Experts compression framework....
This repository contains custom APEX (Adaptive Precision for EXpert Models) -inspired quants, tuned specifically to the underlying mixture-of-experts...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
> A 34%-expert-pruned GLM-5.2 that holds parity with the full unpruned model on a > well-powered real-world eval...
MiniMax-M3-REAP22-Coder A JANG-quantized MiniMax-M3 — coding/agentic + multimodal — for the vMLX engine (Apple Silicon / MLX). >...
This repository now bundles glmmoedsa.py (declared via modelfile in config.json`), a fixed runtime for this architecture, and needs...
static quants of https://huggingface.co/dervig/m51Lab-MiniMax-M2.7-REAP-139B-A10B For a convenient overview and download list, visit our model page for this model....