unsloth/DeepScaleR-1.5B-Preview - Каталог нейросетей
Генерация текста

unsloth/DeepScaleR-1.5B-Preview

Добавлено:
unsloth/DeepScaleR-1.5B-Preview

We have a free Google Colab notebook for turning Llama 3.1 (8B) into a reasoning model: https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Llama3.1_(8B)-GRPO.ipynb All notebooks are beginner friendly! Add your dataset, click «Run All», and you’ll get a 2x faster finetuned model which can be exported to GGUF, vLLM or uploaded to Hugging Face. — This Llama 3.2 conversational notebook-Conversational.ipynb) is useful for ShareGPT ChatML / Vicuna templates. — This text completion notebook-TextCompletion.ipynb) is for raw text. This DPO notebook replicates Zephyr. — Kaggle has 2x T4s, but we use 1. Due to overhead, 1x T4 is 5x faster. DeepScaleR-1.5B-Preview 🚀 Democratizing Reinforcement Learning for LLMs 🌟 DeepScaleR-1.5B-Preview is a language model fine-tuned from DeepSeek-R1-Distilled-Qwen-1.5B using distributed reinforcement learning (RL) to scale up to long context lengths. The model achieves 43.1% Pass@1 accuracy on AIME 2024, representing a 15% improvement over the base model (28.8%) and surpassing OpenAI’s O1-Preview performance with just 1.5B parameters. Our training dataset consists of approximately 40,000 unique problem-answer pairs compiled from: -…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: unsloth
Теги: qwen2, deepseek, unsloth, qwen, conversational, en, text-generation-inference, endpoints_compatible
Лайков: 3  |  Загрузок: 23

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.