This repository contains the Guru-7B (base Qwen2.5-7B) model presented in Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective. The leaderboard is evaluated with our evaluation code. The parameters we set in evaluation for all models: temperature=1.0, top_p=0.7.
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: IFM
Теги: qwen2, conversational, text-generation-inference, endpoints_compatible
Лайков: 4 | Загрузок: 110
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.