nvidia/Llama-3.3-Nemotron-70B-Reward - Каталог нейросетей
Генерация текста

nvidia/Llama-3.3-Nemotron-70B-Reward

Добавлено:
nvidia/Llama-3.3-Nemotron-70B-Reward

Llama-3.3-Nemotron-70B-Reward is a large language model that leverages Meta-Llama-3.3-70B-Instruct as the foundation and is fine-tuned using scaled Bradley-Terry modeling to predict the quality of LLM generated responses. Given an English conversation with multiple turns between user and assistant (of up to 4,096 tokens), it rates the quality of the final assistant turn using a reward score. For the same prompt, a response with higher reward score has higher quality than another response with a lower reward score, but the same cannot be said when comparing the scores between responses to different prompts. As of 15 May 2025, this model achieves the highest JudgeBench at 73.7% and second highest on RM-Bench at 79.9% among Bradley-Terry Reward Models. See details on how this model was trained at https://arxiv.org/abs/2505.11475 GOVERNING TERMS: Use of this model is governed by the NVIDIA Open Model License . Additional Information: Llama 3.3 Community License Agreement. Built with Llama. Llama-3.3-Nemotron-70B-Reward labels an LLM-generated response to a user query with a reward score. HuggingFace 06/27/2025 via https://huggingface.co/nvidia/Llama-3.3-Nemotron-70B-Reward…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: nvidia
Теги: llama, nvidia, llama3.3, conversational, en, text-generation-inference
Лайков: 4  |  Загрузок: 106

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.