This is the repo for a new financial domain large language model, InvestLM, tuned on Mixtral-8x7B-v0.1, using a carefully curated instruction dataset related to financial investment. We provide guidance on how to use InvestLM for inference. AWQ is an efficient, accurate, and blazing-fast low-bit weight quantization method, currently supporting 4-bit quantization. Compared to GPTQ, it offers faster Transformers-based inference. Please use the following command to log in hugging face first.
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: yixuantt
Теги: mixtral, conversational, en, text-generation-inference, endpoints_compatible, 4-bit, awq
Лайков: 3 | Загрузок: 0
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.