GGUF build of pottokao/Ornith-1.5-35B-A3B-abliterated-NVFP4-DFlash, for use with llama.cpp. The 4-bit weights are repacked bit-exact from the NVFP4 checkpoint into GGMLTYPENVFP4 — they are not dequantized and re-quantized, so there is no double-quantization penalty. On Blackwell GPUs llama.cpp runs these through native FP4 tensor cores. 19.5 GB, plus a 772 MB DFlash draft model for speculative decoding. Runs on 2×16 GB consumer GPUs (tested on 2× RTX 5070 Ti). > ⚠️ Text-only. No vision tower, no MTP head (the abliteration was done on a > language-model-only export). Converted with —no-mtp. > ⚠️ Uncensored. Safety refusal behaviour has been deliberately removed. You are responsible > for how you use it. Tooling derived from remove-refusals-with-transformers. BF16 weights: pottokao/Ornith-1.5-35B-A3B-abliterated. NVIDIA TensorRT Model Optimizer 0.45.0, per-layer recipe matched exactly to the official ornith-ai/Ornith-1.5-35B-A3B-NVFP4 (verified tensor-by-tensor: weightscale2 30841, inputscale 130, 291 quantized layers, 0 diff in the language model). Calibration: 64 × 512 tokens from abisee/cnndailymail. Conversion (latest llama.cpp, which has a ModelOpt-aware branch): —no-mtp is…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: pottokao
Теги: gguf, llama.cpp, nvfp4, modelopt, abliterated, uncensored, moe, mamba
Лайков: 4 | Загрузок: 4,476
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.