MohamedRashad/AceGPT-13B-chat-AWQ - Каталог нейросетей
Генерация текста

MohamedRashad/AceGPT-13B-chat-AWQ

Добавлено:
MohamedRashad/AceGPT-13B-chat-AWQ

— Model creator: FreedomIntelligence — Original model: AceGPT 13B Chat This repo contains AWQ model files for FreedomIntelligence’s AceGPT 13B Chat. In my effort of making Arabic LLms Available for consumers with simple GPUs I have Quantized two important models: — AceGPT 13B Chat AWQ (We are Here) — AceGPT 7B Chat AWQ AWQ is an efficient, accurate and blazing-fast low-bit weight quantization method, currently supporting 4-bit quantization. Compared to GPTQ, it offers faster Transformers-based inference with equivalent or better quality compared to the most commonly used GPTQ settings. — Text Generation Webui — using Loader: AutoAWQ — vLLM — Llama and Mistral models only — Hugging Face Text Generation Inference (TGI) — Transformers version 4.35.0 and later, from any code or client that supports Transformers — AutoAWQ — for use from Python code — Requires: Transformers 4.35.0 or later. — Requires: AutoAWQ 0.1.6 or later. Note that if you are using PyTorch 2.0.1, the above AutoAWQ command will automatically upgrade you to PyTorch 2.1.0. If you are using CUDA 11.8 and wish to continue using PyTorch 2.0.1, instead run this command: If you have problems installing AutoAWQ using the…

Модальности:
Генерация текста

Области применения:
Диалог / чат


Задача: Генерация текста
Автор: MohamedRashad
Теги: llama, en, ar, text-generation-inference, 4-bit, awq
Лайков: 3  |  Загрузок: 17

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.