Lego-X/qwen3_5_35b_a3b_ohsdk_200k_rl - Каталог нейросетей
Генерация текста

Lego-X/qwen3_5_35b_a3b_ohsdk_200k_rl

Добавлено:
Lego-X/qwen3_5_35b_a3b_ohsdk_200k_rl

Qwen3.5-35B-A3B trained with online RL inside the unmodified OpenHands SDK harness, at 200K context, on 2,699 real repository issues whose own test suites produce the reward. SWE-bench Verified: 64.0 → 70.4 (+6.4) — no reward model, no reference-patch similarity, no harness rewrite. This checkpoint is the OpenHands SDK production run of Lego-RL (Faithful · Reliable · Observable), the RL component of the LegoX series. The agent solves a real issue in a real repository inside a fresh sandbox, the task’s own verifier suite decides {0, 1}, and the trajectory the harness actually produced — token ids, masks, log-probs and MoE expert routes captured inside the serving path — becomes the gradient step. The scaffold is part of the environment, not the policy. The same weights score very differently depending on which harness runs them, so training under a rewritten control flow optimizes for a deployment you never ship: SWE-bench Verified (%), one shared protocol: temperature 0.7, 200 turns, 200K context. Each Lego-RL column is a separate run trained in that harness from the same starting checkpoint, the same 2,699 tasks and the same 3 epochs. This repository is the OpenHands SDK run…

Модальности:
Генерация текста

Области применения:
Генерация кода Диалог / чат


Задача: Генерация текста
Автор: Lego-X
Теги: qwen3_5_moe, image-text-to-text, code, swe-bench, agentic, coding-agent, reinforcement-learning, gspo
Лайков: 4  |  Загрузок: 717

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.