DeepSeek-R1-Distill-Qwen-7B
LiteRT is Google’s on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch,...
LiteRT is Google’s on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch,...
!MMLU-Pro !GSM8K !GPQA !Uncensored All the capability, none of the refusals. > Uncensored at no cost to quality....
NVFP4 (NVIDIA FP4, weight-only NVFP4A16) quantization of empero-ai/Qwythos-9B-Claude-Mythos-5-1M — a Claude-Mythos/Fable-trace reasoning fine-tune of Qwen3.5-9B (qwen35: a dense...
> No matter your GPU. No matter your RAM. With ~4.5 GB of VRAM or unified memory free,...
Mixed-precision MLX quantization of huihui-ai/Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated, quantized with MLX Smart Quantize (MSQ) — my own sensitivity-based mixed-precision quantization method...
GGUF quantizations of empero-ai/Qwythos-9B-Claude-Mythos-5-1M for llama.cpp, Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes. Qwythos-9B is a...
Added a Jinja chat template so the model can format conversations correctly and work smoothly with mlx-lm chat-style...
> VibeThinker-3B-heretic_decensored is a reasoning-focused language model built on top of WeiboAI/VibeThinker-3B and modified using the Heretic abliteration...
VibeThinker-3B-hereticdecensored Reasoning-focused language model modified using the Heretic abliteration toolkit Abliteration 3B Parameters STEM Reasoning Uncensored VibeThinker-3B-hereticdecensored is...
> [!WARNING] > This model is a beta preview and a research artifact. It is not an open...