Qwen3.6-35B-A3B-NVFP4-MTP-GGUF
RTX 5090 Windows native GGUF comparison — updated July 19, 2026 Native llama.cpp conversions of NVIDIA and Unsloth...
RTX 5090 Windows native GGUF comparison — updated July 19, 2026 Native llama.cpp conversions of NVIDIA and Unsloth...
GGUF conversion of z-lab/Qwen3.6-35B-A3B-DFlash for llama.cpp. > This is a DFlash draft model, not a standalone language model....
> Uncensored, agentic model — tool-use, shell & coding, zero refusals. — 🧭 FableForge Nexus — 🍺 Infinite...
A Hindi instruction-tuned fine-tune of Gemma 4 E4B, quantized to GGUF for local / CPU / edge use...
This repository hosts GGUF weights for Anubis-Mini-8B-v1-heretic, quantized from the source floating-point tensors provided by coder3101/Anubis-Mini-8B-v1-heretic. 🔄 Sister...
Converted to BF16 using converthftogguf.py, then quantized using llama-quantize` from llama.cpp. Did not use an imatrix for any...
> 👋 I’m back. Active development has resumed — expect improvements soon. > ⚠️ Experimental. ASHQ1 is a...
https://huggingface.co/micymike/CodeMate-Qwen-1.5B-32K-Distilled-on-Claude-Fable-5 This project is an independent research effort and is not affiliated with or endorsed by Anthropic, Claude,...
This model was converted to GGUF format from Qwen/Qwen-AgentWorld-35B-A3B using llama.cpp via the ggml.ai’s GGUF-my-repo space. Refer to...
GGUF quantizations of empero-ai/Qwythos-9B-Claude-Mythos-5-1M for llama.cpp, Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes. Qwythos-9B is a...