This model is based on the paper A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning. By applying SFT and GRPO on difficult math problems, we enhanced the performance of DeepSeek-R1-Distill-Qwen-14B and developed Fast-Math-R1-14B, which achieves approx. 30% faster inference on average, while maintaining accuracy. In addition, we trained and open-sourced Fast-OpenMath-Nemotron-14B, an efficiency-optimized version of NVIDIA’s OpenMath-Nemotron-14B, following the same approach. Compared to OpenMath-Nemotron-14B, this model enables approx. 30% faster inference on average, with minimal loss in performance. Technical details can be found in our github repository. Project page: https://analokmaus.github.io/Fast-Math-R1/
Модальности:
Генерация текста
Области применения:
Диалог / чат Математика
Задача: Генерация текста
Автор: RabotniKuma
Теги: qwen2, mathematical-reasoning, conversational, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 22
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.