gaotang/RM-R1-Qwen2.5-Instruct-7B - Каталог нейросетей
Генерация текста

gaotang/RM-R1-Qwen2.5-Instruct-7B

Добавлено:
gaotang/RM-R1-Qwen2.5-Instruct-7B

RM-R1 is a training framework for Reasoning Reward Model (ReasRM) that judges two candidate answers by first thinking out loud—generating structured rubrics or reasoning traces—then emitting its preference. Compared to traditional scalar or generative reward models, RM-R1 delivers state-of-the-art performance on public RM benchmarks on average while offering fully interpretable justifications. Two-stage training 1. Distillation of ~8.7 K high-quality reasoning traces (Chain-of-Rubrics). 2. Reinforcement Learning with Verifiable Rewards** (RLVR) on ~64 K preference pairs. Backbones** released: 7 B / 14 B / 32 B Qwen-2.5-Instruct variants + DeepSeek-distilled checkpoints. RLHF / RLAIF: plug-and-play reward function for policy optimisation. Automated evaluation: LLM-as-a-judge for open-domain QA, chat, and reasoning. Research**: study process supervision, chain-of-thought verification, or rubric generation. Try the model with this example. Full demo notebook available at:

Модальности:
Генерация текста

Области применения:
Диалог / чат Следование инструкциям


Задача: Генерация текста
Автор: gaotang
Теги: qwen2, conversational, en, text-generation-inference, endpoints_compatible
Лайков: 4  |  Загрузок: 229

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.