openbmb/RLPR-Llama3.1-8B-Inst - Каталог нейросетей
Генерация текста

openbmb/RLPR-Llama3.1-8B-Inst

Добавлено:
openbmb/RLPR-Llama3.1-8B-Inst

RLPR-Llama3.1-8B-Inst is trained from Llama3.1-8B-Inst with the RLPR framework, which eliminates reliance on external verifiers and is simple and generalizable for more domains. 💡 Verifier-Free Reasoning Enhancement: RLPR pioneers reinforcement learning for reasoning tasks by leveraging the LLM’s intrinsic generation probability as a direct reward signal. This eliminates the need for external verifiers and specialized fine-tuning, offering broad applicability and effectively handling complex, diverse answers. 🛠️ Innovative Reward & Training Framework: Features a robust Probability-based Reward (PR) using average decoding probabilities of reference answers for higher quality, debiased reward signals, outperforming naive sequence likelihood. Implements an standard deviation filtering mechanism that dynamically filters prompts to stabilize training and significantly boost final performance. 🚀 Strong Performance in General & Mathematical Reasoning:** Demonstrates substantial reasoning improvements across diverse benchmarks, surpassing the RLVR baseline for 1.4 average points across seven benchmarks. — Trained from model: Llama-3.1-8B-Instruct — Trained on data:…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: openbmb
Теги: llama, text-generation-inference, conversational, en, endpoints_compatible
Лайков: 4  |  Загрузок: 96

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.