Chat & support: TheBloke’s Discord server Want to contribute? TheBloke’s Patreon page TheBloke’s LLM work is generously supported by a grant from andreessen horowitz (a16z) These files are GPTQ 4bit model files for OpenAccess AI Collective’s Minotaur 13B Fixed merged with Kaio Ken’s SuperHOT 8K. It is the result of quantising to 4bit using GPTQ-for-LLaMa. This is an experimental new GPTQ which offers up to 8K context size The increased context is tested to work with ExLlama, via the latest release of text-generation-webui. It has also been tested from Python code using AutoGPTQ, and trustremotecode=True. Code credits: — Original concept and code for increasing context length: kaiokendev — Updated Llama modelling code that includes this automatically via trustremotecode: emozilla. GGML versions are not yet provided, as there is not yet support for SuperHOT in llama.cpp. This is being investigated and will hopefully come soon. 4-bit GPTQ models for GPU inference 2, 3, 4, 5, 6 and 8-bit GGML models for CPU inference Unquantised SuperHOT fp16 model in pytorch format, for GPU inference and for further conversions Unquantised base fp16 model in pytorch format, for GPU inference…
Модальности:
Генерация текста
Задача: Генерация текста
Автор: TheBloke
Теги: llama, custom_code, text-generation-inference, 4-bit, gptq
Лайков: 3 | Загрузок: 23
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.