dealignai/Nemotron-3.5-Lightning-30B-A3B-CRACK-GGUF - Каталог нейросетей
Генерация текста

dealignai/Nemotron-3.5-Lightning-30B-A3B-CRACK-GGUF

Добавлено:
dealignai/Nemotron-3.5-Lightning-30B-A3B-CRACK-GGUF

CRACK-abliterated NVIDIA Nemotron 3.5 Lightning 30B-A3B — GGUF quants for llama.cpp. Three quantizations (Q80 / Q4KM / Q2K) in one repository, each with the native MTP (Multi-Token Prediction) draft head folded in for speculative decoding. Refusal behavior removed while preserving the model’s knowledge, reasoning (thinking), and multilingual ability. > Research artifact with reduced safety guardrails. Use responsibly and lawfully. — Base: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B — hybrid Mamba-2 SSM + MoE + attention (52 layers: 23 Mamba-2 / 23 MoE / 6 attention), 128 routed experts (~3B active), 262K context, reasoning (thinking) ON by default, Multi-Token-Prediction head. — CRACK: early decision-zone abliteration with per-layer refusal directions. Knowledge preserved; the native MTP draft head and reasoning are kept intact. All three include the folded native MTP block (blk.52.nextn.) for speculative decoding. MMLU is logit-mode accuracy (base vs. CRACK — measures knowledge retention). HarmBench is answer-channel compliance on harm behaviors, counting only coherent responses (gibberish/degenerate outputs do not count as compliant). Knowledge is largely retained (overall Δ…

Модальности:
Генерация текста

Области применения:
Логика и рассуждение Диалог / чат


Задача: Генерация текста
Автор: dealignai
Теги: gguf, llama.cpp, nemotron-h, mamba2, moe, abliterated, uncensored, crack
Лайков: 4  |  Загрузок: 16,679

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.