spiral-rl/Spiral-Qwen3-4B - Каталог нейросетей
Генерация текста

spiral-rl/Spiral-Qwen3-4B

Добавлено:
spiral-rl/Spiral-Qwen3-4B

This model is trained with self-play on multi-games (TicTacToe, Kuhn Poker, Simple Negotiation) using the SPIRAL framework. Recent advances in reinforcement learning have shown that language models can develop sophisticated reasoning through training on tasks with verifiable rewards, but these approaches depend on expert-curated problem-answer pairs and domain-specific reward engineering. We introduce SPIRAL, a self-play framework where models learn by playing multi-turn, zero-sum games against continuously improving versions of themselves, eliminating the need for human supervision. Through zero-sum self-play, SPIRAL generates an infinite curriculum of progressively challenging problems as models must constantly adapt to stronger opponents. Applying SPIRAL to Qwen3 base models in two-player zero-sum text games, we observe the agents develop advanced reasoning strategies to win the competitive game. Furthermore, the trained models show substantial gains on a range of math and general reasoning benchmarks. These results suggest that self-play in zero-sum games can naturally induce transferable reasoning capabilities, highlighting a promising direction for autonomous reasoning…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: spiral-rl
Теги: qwen3, conversational, text-generation-inference, endpoints_compatible
Лайков: 4  |  Загрузок: 21

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.