qwen3_5_35b_a3b_ohsdk_200k_rl
Qwen3.5-35B-A3B trained with online RL inside the unmodified OpenHands SDK harness, at 200K context, on 2,699 real repository...
Qwen3.5-35B-A3B trained with online RL inside the unmodified OpenHands SDK harness, at 200K context, on 2,699 real repository...
An RL-trained agent whose job is to write RL training jobs for smaller models. This is the LoRA...
Hiro-Pharma is a causal language model based on Qwen/Qwen2.5-7B. It is released in safetensors format under the Apache...
We present GoLongRL, a fully open-source, capability-oriented post-training recipe for long-context reinforcement learning with verifiable rewards (RLVR). Overall...
ToolOmni is a tool-use language model released for the ACL 2026 Main Conference paper ToolOmni: Enabling Open-World Tool...
Strongest CodeScout model — open-source SOTA on SWE-Bench code localization. CodeScout-14B is part of the CodeScout family of...
Starting from nothing but 9 search queries, we used the Lightning Rod SDK to automatically generate 3,178 forecasting...
Starting from nothing but 5 search queries, we used the Lightning Rod SDK to automatically generate 2,108 forecasting...
hkust-nlp/drkernel-8b is a Qwen3-8B-based model specialized for GPU kernel generation and optimization (especially Triton) in the DR.Kernel framework....
Maincoder-1B-ONNX — это ONNX-оптимизированная версия Maincoder-1B, кодо-ориентированной языковой модели, оптимизированной для задач генерации и завершения кода. Эта версия...