> VibeThinker-3B-heretic_decensored is a reasoning-focused language model built on top of WeiboAI/VibeThinker-3B and modified using the Heretic abliteration toolkit. The model applies refusal-direction analysis and targeted weight-space interventions to reduce internal refusal behaviors while preserving the strong mathematical, coding, and STEM reasoning capabilities inherited from the VibeThinker training pipeline. > About VibeThinker-3B: VibeThinker-3B is a 3-billion-parameter reasoning-focused language model developed by WeiboAI. Built on top of Qwen2.5-Coder-3B, it was trained using the Spectrum-to-Signal Principle (SSP) post-training pipeline, combining curriculum-based two-stage supervised fine-tuning, multi-domain reinforcement learning through MaxEnt-Guided Policy Optimization (MGPO), offline self-distillation, and instruction-following reinforcement learning. > [!IMPORTANT] > This model is intended strictly for research and learning purposes. Due to reduced internal refusal mechanisms, it may generate sensitive or unrestricted content. Users assume full responsibility for how the model is used. The authors and hosting platform disclaim any liability for generated outputs.…
Модальности:
Генерация текста
Области применения:
Генерация кода Математика Логика и рассуждение Диалог / чат
Задача: Генерация текста
Автор: prithivMLmods
Теги: gguf, llama-cpp, text-generation-inference, math, code, reasoning, gpqa, instruction-following
Лайков: 4 | Загрузок: 477
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.