Qwen3.8-27B-DFlash2-Q3_K_M-GGUF
Two quantizations of the DFlash 2 draft model for Qwen/Qwen3.8-27B, both calibrated with an importance matrix captured from...
Two quantizations of the DFlash 2 draft model for Qwen/Qwen3.8-27B, both calibrated with an importance matrix captured from...
same grug as ProCreations/grug-v1.1-qwen-3.8-27b, plus a draft head that guess ahead, tuned on grug own output. Qwen3.8 ship...
A reproducible deployment and benchmark package for the official deepseek-ai/DeepSeek-V4-Flash-0731 checkpoint on two NVIDIA GB10-class nodes using vLLM...
A 4-bit ROCmFP4STRIX` quant of poolside/Laguna-S-2.1 (118B-A8B), built and measured on an AMD Ryzen AI Max+ 395 (Strix...
A 6-bit MLX build of the Fable-Fusion-711 tune with one directional refusal ablation baked into the weights. The...
This repository contains a DFlash draft model for moonshotai/Kimi-K3 trained only on a generic data mix (no tool...
A W4A16-quantized DSpark speculator for zai-org/GLM-5.2-FP8, tuned for memory-bandwidth-bound hardware — specifically a two-node NVIDIA DGX Spark (GB10)...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
RTX 5090 Windows native GGUF comparison — updated July 19, 2026 Native llama.cpp conversions of NVIDIA and Unsloth...