Llama-3.1-Nemotron-70B-Instruct-AWQ-INT4
If you run into errors on a multi GPU machine, I’ve found that setting CUDAVISIBLEDEVICES=0 helps. Llama-3.1-Nemotron-70B-Instruct is...
If you run into errors on a multi GPU machine, I’ve found that setting CUDAVISIBLEDEVICES=0 helps. Llama-3.1-Nemotron-70B-Instruct is...
Преобразованная и квантованная контрольная точка nvidia/Nemotron-4-340B-Instruct. В частности, он был получен из контрольной точки v1.0 .nemo на NGC....
FP8 квантованная контрольная точка nvidia/Minitron-8B-Base для использования с vLLM. Модальности:Генерация текста Задача: Генерация текста Автор: mgoin Теги: nemotron,...
FP8 квантованная контрольная точка nvidia/Minitron-4B-Base для использования с vLLM. Модальности:Генерация текста Задача: Генерация текста Автор: mgoin Теги: nemotron,...
> Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no...
Community GGUF quantizations of nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16. These files contain the standard 52-layer inference model. The trailing MTP prediction head...
Luciole-23B-Instruct-1.1-FP8 is a FP8-quantized version of Luciole-23B-Instruct-1.1 in the transformers / compressed-tensors (https://github.com/neuralmagic/compressed-tensors) format, quantized with LLM Compressor...
Model Description Bias, Risks, and Limitations Recommendations Training Details Training Data Instruction template Training Procedure Evaluation Testing the...
QLoRA fine-tune of NVIDIA Nemotron-3-Super-120B-A12B (120B-total / 12B-active hybrid Mamba-2 + Latent-MoE, nemotronh) on the piagent split of...
QLoRA fine-tune of NVIDIA Nemotron-3-Super-120B-A12B (120B-total / 12B-active hybrid Mamba-2 + Latent-MoE, nemotronh) on the piagent split of...