Runs smoothly on single 3090 in webui with context length set to 4096, ExLlamav2HF loader and cache8bit=True All comments are greatly appreciated, download, test and if you appreciate my work, consider buying me my fuel:
Модальности:
Генерация текста
Задача: Генерация текста
Автор: TeeZee
Теги: llama, merge, not-for-all-audiences, text-generation-inference, endpoints_compatible
Лайков: 3 | Загрузок: 8
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.