sahilchachra/MiniCPM5-1B-Uncensored - Каталог нейросетей
Генерация текста

sahilchachra/MiniCPM5-1B-Uncensored

Добавлено:
sahilchachra/MiniCPM5-1B-Uncensored

A fully uncensored version of openbmb/MiniCPM5-1B, produced with a single training-free stage: single-direction abliteration (Arditi et al., 2024). Refusals on AdvBench drop from 85% → 2% with zero over-refusal regression on benign prompts — no fine-tuning, no new data, weights edited directly. > Intended for: security research, red-teaming, jailbreak benchmarking, and AI-safety study. Not intended for production deployment or harmful use. A 83-point drop in harmful refusals while preserving benign behavior. — Abliteration is surgical, not lossless — removing the refusal direction can occasionally affect responses that legitimately overlap with it. General reasoning and benign behavior are preserved (0% over-refusal on the benign set). — No new knowledge — abliteration only removes refusal behavior; it adds no information or capability. — Small model — at ~1B parameters, factual accuracy and complex reasoning are limited regardless of alignment. — Responsible use — published for safety research and red-teaming. The authors do not endorse harmful use of this model.

Модальности:
Генерация текста

Области применения:
Логика и рассуждение Диалог / чат


Задача: Генерация текста
Автор: sahilchachra
Теги: mlx, llama, uncensored, abliteration, safety-research, reasoning, minicpm, conversational
Лайков: 4  |  Загрузок: 533

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.