Ornith-1.5-35B-A3B-abliterated-NVFP4-DFlash-GGUF
GGUF build of pottokao/Ornith-1.5-35B-A3B-abliterated-NVFP4-DFlash, for use with llama.cpp. The 4-bit weights are repacked bit-exact from the NVFP4 checkpoint...
GGUF build of pottokao/Ornith-1.5-35B-A3B-abliterated-NVFP4-DFlash, for use with llama.cpp. The 4-bit weights are repacked bit-exact from the NVFP4 checkpoint...
NVFP4 (W4A16) quantization of an abliterated (refusal-direction removed) build of ornith-ai/Ornith-1.5-35B-A3B. The per-layer quantization recipe is matched exactly...
> 短思考版 Qwen3.8-27B —— 保留推理质量,压缩思考长度 > Short-thinking Qwen3.8-27B — keep reasoning quality, compress thinking length ShortThink is a...
A second quantization of inclusionAI/Ling-3.0-flash, alongside Ling-3.0-flash-HybridQuant-NVFP4-W4A16. Identical bit placement and identical serving contract; it differs in how...
A.X K2 — это крупномасштабная языковая модель Mixure-of-Experts (MoE), обученная с нуля как высокопроизводительная агентная базовая модель и...
This is a full-expert GLM-5.2 hybrid checkpoint that combines the compact MXFP8/NVFP4/NF3 layout from madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid with selected refusal-reduction...
Non-uniform expert prune · NVFP4 · runs on FOUR Blackwell cards · verified 0-refusal Three transformations in one...
This is a reasoning-first, agentic and conversational merge of google/gemma-4-31B-it. It has enough creative and roleplay tuning to...
Калиброванное квантование NVFP4 InternScience/Agents-A1 (агент Qwen3.5-35B-A3B hybrid MoE) для vLLM. 21,8 ГБ — и оно соответствует или превосходит...
RTX 5090 Windows native GGUF comparison — updated July 19, 2026 Native llama.cpp conversions of NVIDIA and Unsloth...