Chat & support: TheBloke’s Discord server Want to contribute? TheBloke’s Patreon page TheBloke’s LLM work is generously supported by a grant from andreessen horowitz (a16z) These files were quantised using hardware kindly provided by Massed Compute. AWQ is an efficient, accurate and blazing-fast low-bit weight quantization method, currently supporting 4-bit quantization. Compared to GPTQ, it offers faster Transformers-based inference with equivalent or better quality compared to the most commonly used GPTQ settings. AWQ models are currently supported on Linux and Windows, with NVidia GPUs only. macOS users: please use GGUF models instead. — Text Generation Webui — using Loader: AutoAWQ — vLLM — version 0.2.2 or later for support for all model types. — Hugging Face Text Generation Inference (TGI) — Transformers version 4.35.0 and later, from any code or client that supports Transformers — AutoAWQ — for use from Python code AWQ model(s) for GPU inference. GPTQ models for GPU inference, with multiple quantisation parameter options. 2, 3, 4, 5, 6 and 8-bit GGUF models for CPU+GPU inference Undi’s original unquantised fp16 model in pytorch format, for GPU inference and for…
Модальности:
Генерация текста
Задача: Генерация текста
Автор: TheBloke
Теги: mistral, not-for-all-audiences, nsfw, text-generation-inference, 4-bit, awq
Лайков: 3 | Загрузок: 13
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.