Nickyang/FastCuRL-1.5B-V3 - Каталог нейросетей
Генерация текста

Nickyang/FastCuRL-1.5B-V3

Добавлено:
Nickyang/FastCuRL-1.5B-V3

We release FastCuRL-1.5B-Preview, a slow-thinking reasoning model that outperforms the previous SoTA DeepScaleR-1.5B-Preview with 50% training steps! We adapt a novel curriculum-guided iterative lengthening reinforcement learning to the DeepSeek-R1-Distill-Qwen-1.5B and observe continuous performance improvement as training steps increase. To better reproduce our work and advance research progress, we open-source our code, model, and data. We report Pass@1 accuracy averaged over 16 samples for each problem. Following DeepScaleR, our training dataset consists of 40,315 unique problem-answer pairs compiled from: — AIME problems (1984-2023) — AMC problems (before 2023) — Omni-MATH dataset — Still dataset — Our training experiments are powered by our heavily modified fork of verl and deepscaler. — Our model is trained on top of DeepSeek-R1-Distill-Qwen-1.5B.

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: Nickyang
Теги: qwen2, conversational, en, text-generation-inference, endpoints_compatible
Лайков: 4  |  Загрузок: 25

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.