taardis-27b-full-ternary
Тройное адаптивное выравнивание И V2 поставляет систему коррекции трубопровода: The Doctors — 496 кросс-слойных тройных ветвей низкого ранга,...
Тройное адаптивное выравнивание И V2 поставляет систему коррекции трубопровода: The Doctors — 496 кросс-слойных тройных ветвей низкого ранга,...
Q8_0 квантования XHToken/Spark-X2.5-4B, преобразованного из официального BF16 GGUF с совместимой реализацией XHToken/llama.cpp. Для поддержки Spark-X2.5 требуется совместимая реализация...
GGUF quantizations of Akicou/Qwen3.8-Flash-Next-REAM-60Pct, the REAM-compressed (Merged) version of Qwen/Qwen3.8-Flash-Next. REAM (Router Expert Activation Merging) pruned 40% of...
This experimental text-only build compresses Qwen3.8-Flash-Next from a 72.55 GB calibrated source GGUF to 39,721,239,200 bytes: — 36.993287...
GGUF quantizations of GestaltLabs/Ornstein3.8-27B: a Qwen 3.8 27B dense vision-language fine-tune with interleaved linear and full attention (Gated...
GGUF quantizations of zai-org/GLM-5.3-Flash, made with llama.cpp. 320B total parameters, 18B active. 45 layers with hybrid attention: 34...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
Two quantizations of the DFlash 2 draft model for Qwen/Qwen3.8-27B, both calibrated with an importance matrix captured from...
GGUF build of pottokao/Ornith-1.5-35B-A3B-abliterated-NVFP4-DFlash, for use with llama.cpp. The 4-bit weights are repacked bit-exact from the NVFP4 checkpoint...
> 短思考版 Qwen3.8-27B —— 保留推理质量,压缩思考长度 > Short-thinking Qwen3.8-27B — keep reasoning quality, compress thinking length ShortThink is a...