QLoRA fine-tune of NVIDIA Nemotron-3-Super-120B-A12B (120B-total / 12B-active hybrid Mamba-2 + Latent-MoE, nemotronh) on the piagent split of Glint-Research/Fable-5-traces, targeting reasoning, agentic planning, and tool-use. This repo holds the full-precision BF16 weights (the LoRA delta merged into the base, ≈231 GB across 50 shards), loadable directly with 🤗 Transformers. It is the reference master from which the GGUF / NVFP4 / MLX exports derive. INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED BUT SOME MAY HAVE NO CONTENT > These are full-precision Transformers weights, not an LM Studio format (use the > MLX-4bit or > GGUF siblings for LM > Studio) — but the empty init.py that ships with nemotronh repos triggers that bug for > anyone pulling this repo through LM Studio’s downloader. As a precaution it now ships a > non-empty init.py` so such downloads finalize normally. — GGUF Q4KM (LM Studio / llama.cpp): greghavens/fabletron-nemotron-3-super-120b-GGUF — NVFP4 (vLLM / Blackwell FP4):…
Модальности:
Генерация текста
Области применения:
Логика и рассуждение Диалог / чат Вызов функций (Tool use)
Задача: Генерация текста
Автор: greghavens
Теги: nemotron_h, nemotron, mamba2, mixture-of-experts, qlora, reasoning, agentic, tool-use
Лайков: 4 | Загрузок: 84
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.