Vinnnf/Thinkless-1.5B-RL-DeepScaleR - Каталог нейросетей
Генерация текста

Vinnnf/Thinkless-1.5B-RL-DeepScaleR

Добавлено:
Vinnnf/Thinkless-1.5B-RL-DeepScaleR

📄 Paper Link ArXiv 💻 RL Code VainF/Thinkless 💻 SFT Code VainF/Reasoning-SFT 🤖 RL Model Thinkless-1.5B-RL-DeepScaleR 🐣 Warmup Model Thinkless-1.5B-Warmup 📊 Data for Warmup Hybrid-OpenThoughts2-1M-1.5B 📊 Data for RL agentica-org/DeepScaleR-Preview-Dataset We propose Thinkless, a learnable framework that empowers an LLM to adaptively select between short-form and long-form reasoning, based on both task complexity and the model’s ability. Thinkless is trained under a reinforcement learning paradigm and employs two control tokens, for concise responses and for detailed reasoning. At the core of our method is a Decoupled Group Relative Policy Optimization (DeGRPO) algorithm, which decomposes the learning objective of hybrid reasoning into two components: (1) a control token loss that governs the selection of the reasoning mode, and (2) a response loss that improves the accuracy of the generated answers. This decoupled formulation enables fine-grained control over the contributions of each objective, stabilizing training and effectively preventing collapse observed in vanilla GRPO. Empirically, on several benchmarks such as Minerva Algebra, MATH-500, and GSM8K, Thinkless…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: Vinnnf
Теги: qwen2, conversational, text-generation-inference, endpoints_compatible
Лайков: 4  |  Загрузок: 138

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.