NVFP4 (W4A4) quantization of huihui-ai/Huihui-LFM2.5-8B-A1B-abliterated — the abliterated (refusal-reduced) build of Liquid AI’s 8.3B-total / 1.5B-active mixture-of-experts reasoner (131 K context). Runs on one 16 GB Blackwell GPU with room for a ~1.1 M-token KV cache (fp8), and scales cleanly to 2–4 GPUs. Quantized by Lna-Lab with NVIDIA TensorRT Model-Optimizer (modelopt). To our knowledge this is the first NVFP4 build of an abliterated lfm2moe`. > ⚠️ Uncensored model. The base is abliterated by huihui.ai — its safety filtering is significantly reduced. See Usage warnings below. (Per the base card, the abliteration touches the dense path; the MoE experts were not ablated.) > Why it’s nice: 8.3B total / 1.5B active MoE + a hybrid backbone (only 6 of 24 layers > are attention; the rest are short-convolution) means the KV cache is tiny. Shrink the > weights to 4-bit and the freed VRAM turns straight into concurrency — one card > happily serves a stack of parallel sessions, and TP fans that out further. — Single-stream scales with TP: 130 → 210 → 305 tok/s (1→2→4 GPU). — Aggregate scales near-linearly with concurrency; TP=4 reaches ~4.5 K tok/s at C=32 and still has KV headroom for…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: sakamakismile
Теги: vllm, lfm2_moe, nvfp4, fp4, modelopt, moe, quantized, blackwell
Лайков: 4 | Загрузок: 291
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.