TheBloke/Airoboros-L2-70B-3.1.2-AWQ - Каталог нейросетей
Генерация текста

TheBloke/Airoboros-L2-70B-3.1.2-AWQ

Добавлено:
TheBloke/Airoboros-L2-70B-3.1.2-AWQ

Chat & support: TheBloke’s Discord server Want to contribute? TheBloke’s Patreon page TheBloke’s LLM work is generously supported by a grant from andreessen horowitz (a16z) — Model creator: Jon Durbin — Original model: Airoboros L2 70B 3.1.2 This repo contains AWQ model files for Jon Durbin’s Airoboros L2 70B 3.1.2. AWQ is an efficient, accurate and blazing-fast low-bit weight quantization method, currently supporting 4-bit quantization. Compared to GPTQ, it offers faster Transformers-based inference with equivalent or better quality compared to the most commonly used GPTQ settings. — Text Generation Webui — using Loader: AutoAWQ — vLLM — Llama and Mistral models only — Hugging Face Text Generation Inference (TGI) — AutoAWQ — for use from Python code AWQ model(s) for GPU inference. GPTQ models for GPU inference, with multiple quantisation parameter options. 2, 3, 4, 5, 6 and 8-bit GGUF models for CPU+GPU inference Jon Durbin’s original unquantised fp16 model in pytorch format, for GPU inference and for further conversions For my first release of AWQ models, I am releasing 128g models only. I will consider adding 32g as well if there is interest, and once I have done…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: TheBloke
Теги: llama, conversational, text-generation-inference, 4-bit, awq
Лайков: 3  |  Загрузок: 24

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.