This model is a fine-tuned version of mistralai/Mistral-Large-Instruct-2407 trained on long-form, chain-of-thought responses generated using the Qwen/QwQ-32B-Preview model. Mistral-Large-2407-LongCoT is specifically trained to provide detailed, step-by-step reasoning for complex questions. The model prioritizes openness and transparent thinking processes in its responses. The model was fine-tuned on the long-cot-4k dataset, which contains: — Question→Answer pairs — Response lengths ranging from 1024 to 4096 tokens — Step-by-step reasoning patterns — Detailed explanations and calculations when applicable The model uses ChatML format. Here’s an example of proper formatting: — Long-form responses (1024-4096 tokens) — Step-by-step reasoning — Detailed explanations — Mathematical calculations when relevant — Transparent thought processes Here’s a sample conversation demonstrating the model’s capabilities: System: You are a helpful assistant. You should think step-by-step. You should answer all questions by prioritizing the principle of openness. Human: What would happen to the Earth if Saturn suddenly de-spawned? Assistant: So I’ve been thinking about this question: what would happen…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: gghfez
Теги: mistral, text-generation-inference, unsloth, trl, conversational, en, endpoints_compatible
Лайков: 3 | Загрузок: 11
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.