gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF
> No matter your GPU. No matter your RAM. With ~4.5 GB of VRAM or unified memory free,...
> No matter your GPU. No matter your RAM. With ~4.5 GB of VRAM or unified memory free,...
Mixed-precision MLX quantization of huihui-ai/Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated, quantized with MLX Smart Quantize (MSQ) — my own sensitivity-based mixed-precision quantization method...
Added a Jinja chat template so the model can format conversations correctly and work smoothly with mlx-lm chat-style...
> Uncensored, agentic model — tool-use, shell & coding, zero refusals. Base model: Qwen/Qwen3-8B · Part of the...
> Uncensored, agentic model — tool-use, shell & coding, zero refusals. Base model: Qwen/Qwen3-8B · Part of the...
> TL;DR — A local Python-coding assistant that thinks before it codes. 8.25 GB, runs on one 16...
First-ever GGUF conversion of poolside/Laguna-XS.2, produced as part of the Poolside Research Hackathon. Laguna XS.2 is a 33.4B...
Supertron1-8B is an instruction-tuned language model built on top of Qwen3-8B-Base. Designed to be a reliable, efficient daily...
This model was converted to GGUF format from Surpem/Supertron1-4B using llama.cpp. Refer to the original model card for...
With this LLM, we wanted to see, how well tiny LLMs with just 124 million parameters can perform...