We argue that efficient agentic reasoning benefits from decomposing deliberation into three interacting systems: reactive execution (System I) for fine-grained reasoning and direct action; simulative reasoning (System II) that predicts consequences of proposed actions through a world model; and self-regulation (System III) that decides when and how deeply to plan through a learned configurator. SR²AM (Self-Regulated Simulative Reasoning Agentic LLM) is our instantiation: the configurator and simulative planner are realized as distinct stages within an LLM’s chain-of-thought reasoning, with the LLM itself serving as the world model in language space. SR²AM-v0.1-8B achieves an overall Pass@1 of 57.0 across 11 benchmarks spanning math, science, tabular analysis, and web information seeking — competitive with systems at 120–355B parameters. — System I + II + III decomposition: a configurator (System III) decides per-turn whether to plan, continue an existing plan, or act directly; a simulative planner (System II) constructs plans grounded in predicted future states; reactive execution (System I) handles fine-grained reasoning and tool use. — SFT + RL training: supervised learning on…
Модальности:
Генерация текста
Области применения:
Логика и рассуждение Диалог / чат Вызов функций (Tool use)
Задача: Генерация текста
Автор: sailing-lab
Теги: qwen3, agent, reasoning, tool-use, simulative-planning, conversational, en, text-generation-inference
Лайков: 4 | Загрузок: 33
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.