This repository contains nvidia/Llama-3.1-Nemotron-70B-Instruct-HF quantized from 16-bit floats to 4-bit integers, using xMAD.ai proprietary technology. 1. Accuracy: This xMADified model is the best quantized version of the nvidia/Llama-3.1-Nemotron-70B-Instruct-HF model (40 GB only). See Table 1 below for model quality benchmarks. 2. Memory-efficiency: The full-precision model is around 140 GB, while this xMADified model is under 40 GB, making it feasible to run on a single 48 GB GPU. 3. Fine-tuning: These models are fine-tunable over the same reduced (a single 48 GB GPU) hardware in mere 3-clicks. Watch our product demo here Loading the model checkpoint of this xMADified model requires around 40 GB of VRAM. Hence it can be efficiently run on a single 48 GB GPU. 1. Run the following *commands to install the required packages. If you found this model useful, please cite our research paper. For additional xMADified models, access to fine-tuning, and general questions, please contact us at support@xmad.ai and join our waiting list.
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: xmadai
Теги: llama, conversational, endpoints_compatible, 4-bit, gptq
Лайков: 4 | Загрузок: 9
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.