ucalyptus/prem-1B-grpo - Каталог нейросетей
Генерация текста

ucalyptus/prem-1B-grpo

Добавлено:
ucalyptus/prem-1B-grpo

This model was fine-tuned using GRPO (Generative Reward-Powered Optimization), a reinforcement learning technique that optimizes language models using multiple reward functions to improve both output format consistency and mathematical reasoning abilities. — Started with the premai-io/prem-1B-chat base model — Uses Flash Attention 2 for efficient training — Model architecture: Causal Language Model (CLM) GRPO training involves optimizing the model using two specific reward functions: 1. Format Reward Function — Ensures responses follow a strict XML-style format: — Rewards are given for: — Strict format adherence (0.5 points) — Soft format matching (0.3 points) — Integer answer format (0.5 points) — Proper XML structure (up to 0.5 points) 2. Correctness Reward Function — Evaluates mathematical accuracy — Awards 2.0 points for correct numerical answers — Awards 0.0 points for incorrect answers — Learning rate: 5e-6 — Batch size: 2 per device — Gradient accumulation steps: 2 — Training epochs: 1 — Uses cosine learning rate scheduler — Warmup ratio: 0.1 — Uses bfloat16 precision — Maximum prompt length: 256 tokens — Maximum completion length: 200 tokens — Number of generations per…

Модальности:
Генерация текста

Области применения:
Математика Логика и рассуждение Диалог / чат


Задача: Генерация текста
Автор: ucalyptus
Теги: llama, math, reasoning, grpo, gsm8k, reinforcement-learning, conversational, en
Лайков: 3  |  Загрузок: 5

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.