greghavens/fabletron-nemotron-3-super-120b-MLX-4bit - Каталог нейросетей
Генерация текста

greghavens/fabletron-nemotron-3-super-120b-MLX-4bit

Добавлено:
greghavens/fabletron-nemotron-3-super-120b-MLX-4bit

QLoRA fine-tune of NVIDIA Nemotron-3-Super-120B-A12B (120B-total / 12B-active hybrid Mamba-2 + Latent-MoE, nemotronh) on the piagent split of Glint-Research/Fable-5-traces, targeting reasoning, agentic planning, and tool-use. This repo holds the MLX 4-bit export (≈4.5 bits-per-weight, ≈64 GB) for Apple Silicon (Metal) via mlx-lm. > Support caveat: this is a structurally valid MLX 4-bit export (converted and lazy-load > verified). nemotronh is a hybrid Mamba-2 + MoE architecture; confirm your mlx-lm version > actually implements it before relying on this for inference. For the most portable runnable > artifacts, prefer the GGUF Q4KM or NVFP4** siblings. > How it was built: converted with mlx-lm on Linux, CPU-only (pip install mlx mlx-cpu > mlx-lm — the mlx-cpu backend package lets conversion run without a Mac/Metal GPU). It runs > on Apple Silicon. INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED BUT SOME MAY HAVE NO CONTENT

Модальности:
Генерация текста

Области применения:
Логика и рассуждение Диалог / чат Вызов функций (Tool use)


Задача: Генерация текста
Автор: greghavens
Теги: mlx, nemotron_h, nemotron, mamba2, mixture-of-experts, apple-silicon, 4-bit, qlora
Лайков: 4  |  Загрузок: 276

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.