This is quantized version of nvidia/NVIDIA-Nemotron-Nano-9B-v2 created using llama.cpp NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model’s reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in accuracy for harder prompts that require reasoning. Conversely, allowing the model to generate reasoning traces first generally results in higher-quality final solutions to queries and tasks. The model uses a hybrid architecture consisting primarily of Mamba-2 and MLP layers combined with just four Attention layers. For the architecture, please refer to the Nemotron-H tech report. The model was trained using Megatron-LM and NeMo-RL. The supported languages include: English, German, Spanish, French, Italian, and Japanese. Improved using Qwen. GOVERNING TERMS: This trial service is…
Модальности:
Генерация текста
Задача: Генерация текста
Автор: QuantFactory
Теги: gguf, nvidia, en, es, fr, de, it, ja
Лайков: 4 | Загрузок: 184
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.