NVIDIA-Nemotron-Labs-Teacher-Instruction-Following
The post-training data has a cutoff date of May 2026. The pre-training data has a cutoff date of...
The post-training data has a cutoff date of May 2026. The pre-training data has a cutoff date of...
> Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no...
The pre-training data has a cutoff date of September 2025. The post-training data has a cutoff date of...
QLoRA fine-tune of NVIDIA Nemotron-3-Super-120B-A12B (120B-total / 12B-active hybrid Mamba-2 + Latent-MoE, nemotronh) on the piagent split of...
QLoRA fine-tune of NVIDIA Nemotron-3-Super-120B-A12B (120B-total / 12B-active hybrid Mamba-2 + Latent-MoE, nemotronh) on the piagent split of...
An MLX conversion of nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4, quantized for stock mlx-lm: — routed MoE experts: affine int4, group size 32,...
Квантование MLX nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 для Apple Silicon. — Apple Silicon Mac с унифицированной памятью 128 ГБ — mlx-lm >=...
Nemotron-Cascade-2-30B-A3B ships with custom model code (modelingnemotronh.py) that has two bugs exposed by causalconv1d 1.6.x. These are bugs...
Abliterated version of openNemo-9B with safety refusals removed. Built using Snakehead — Empero AI’s internal abliteration tool specialized...
All evaluations were done using NeMo-Skills. We published a tutorial with all details necessary to reproduce our evaluation...