[Megatron-Opus-14B-2.1 ] Exp finetuned from Microsoft’s Phi-4 is a state-of-the-art open model developed with a focus on responsible problem solving and advanced reasoning capabilities. Built upon a diverse blend of synthetic datasets, carefully filtered public domain websites, and high-quality academic books and Q&A datasets, Megatron-Opus-14B-2.1 ensures that small, capable models are trained with datasets of exceptional depth and precision. Megatron-Opus-14B-2.1 adopts a robust safety post-training approach using open-source and in-house synthetic datasets. This involves a combination of SFT (Supervised Fine-Tuning) and iterative DPO (Direct Preference Optimization) techniques, ensuring helpful and harmless outputs across various safety categories. Megatron-Opus-14B-2.1 is fine-tuned on a carefully curated synthetic dataset generated using an advanced pipeline optimized for Chain of Thought (CoT) reasoning and Responsible Problem Breakdown (RPB) methodologies. This ensures that the model excels at: — Logical reasoning — Step-by-step problem-solving — Breaking down complex tasks into manageable parts The dataset also emphasizes responsible decision-making and fairness in…
Модальности:
Генерация текста
Области применения:
Математика Диалог / чат
Задача: Генерация текста
Автор: prithivMLmods
Теги: llama, text-generation-inference, math, phi4, trl, sft, LlamaForCausalLM, conversational
Лайков: 4 | Загрузок: 18
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.