Chat & support: TheBloke’s Discord server Want to contribute? TheBloke’s Patreon page TheBloke’s LLM work is generously supported by a grant from andreessen horowitz (a16z) This repo contains GPTQ model files for Qwen’s Qwen 7B Chat. Multiple GPTQ parameter permutations are provided; see Provided Files below for details of the options provided, their parameters, and the software used to create them. These files were quantised using hardware kindly provided by Massed Compute. GPTQ models for GPU inference, with multiple quantisation parameter options. Qwen’s original unquantised fp16 model in pytorch format, for GPU inference and for further conversions These GPTQ models are known to work in the following inference servers/webuis. — text-generation-webui — KoboldAI United — LoLLMS Web UI — Hugging Face Text Generation Inference (TGI) This may not be a complete list; if you know of others, please let me know! Multiple quantisation parameters are provided, to allow you to choose the best one for your hardware and requirements. Each separate quant is in a different branch. See below for instructions on fetching from different branches. Most GPTQ files are made with AutoGPTQ.…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: TheBloke
Теги: qwen, custom_code, zh, en, 4-bit, gptq
Лайков: 3 | Загрузок: 500
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.