Ornith-1.5-35B-A3B-abliterated-NVFP4-DFlash-GGUF
GGUF build of pottokao/Ornith-1.5-35B-A3B-abliterated-NVFP4-DFlash, for use with llama.cpp. The 4-bit weights are repacked bit-exact from the NVFP4 checkpoint...
GGUF build of pottokao/Ornith-1.5-35B-A3B-abliterated-NVFP4-DFlash, for use with llama.cpp. The 4-bit weights are repacked bit-exact from the NVFP4 checkpoint...
NVFP4 (W4A16) quantization of an abliterated (refusal-direction removed) build of ornith-ai/Ornith-1.5-35B-A3B. The per-layer quantization recipe is matched exactly...
A second quantization of inclusionAI/Ling-3.0-flash, alongside Ling-3.0-flash-HybridQuant-NVFP4-W4A16. Identical bit placement and identical serving contract; it differs in how...
Этот репозиторий содержит черновик модели Gemma 4 31B Instruction-Tuned Assistant, квантованный до собственной точности FP4 (NVFP4) для высокоэффективного...
Этот репозиторий содержит настраиваемую по инструкциям модель Gemma 4 31B, квантованную с нативной точностью FP4 (NVFP4) для высокоэффективного...
NVFP4 (W4A4) quantization of huihui-ai/Huihui-LFM2.5-8B-A1B-abliterated — the abliterated (refusal-reduced) build of Liquid AI’s 8.3B-total / 1.5B-active mixture-of-experts reasoner...
I’ve been trying to get as much performance as possible out of a single DGX Spark with Gemma...
FP8 quantized version of google/gemma-4-E4B-it (8B params, edge model). Produced and maintained by vrfai. This model was quantized...
Trinity-Large-Thinking is a reasoning-optimized variant of Arcee AI’s Trinity-Large family — a 398B-parameter sparse Mixture-of-Experts (MoE) model with...
This repository presents post FP8 quantized Falcon-H1R-7B-FP8 via NVIDIA Model Optimizer, enabling efficient inference while preserving the strong...