This model is super fine-tune with philosophy of science, math, epistemology dataset, to provide high quality responses(from first fine-tune) than Llama-3.1-8B and Google Gemma 2 9B. Super fine tuned with various datasets. The Heavy fine-tuned Mistral-Nemo-Base-2407 Large Language Model (LLM) is a pretrained generative text model of 12B parameters trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size. For more details about this model please refer to our release blog post. — Released under the Apache 2 License — Pre-trained and instructed versions — Trained with a 128k context window — Trained on a large proportion of multilingual and code data — Drop-in replacement of Mistral 7B Mistral Nemo is a transformer model, with the following architecture choices: — Layers: 40 — Dim: 5,120 — Head dim: 128 — Hidden dim: 14,436 — Activation Function: SwiGLU — Number of heads: 32 — Number of kv-heads: 8 (GQA) — Vocabulary size: 217 ~= 128k — Rotary embedd
Модальности:
Генерация текста
Задача: Генерация текста
Автор: EpistemeAI
Теги: mistral, text-generation-inference, unsloth, trl, en, endpoints_compatible
Лайков: 3 | Загрузок: 25
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.