solidrust/Hermes-3-Llama-3.1-8B-AWQ - Каталог нейросетей
Генерация текста

solidrust/Hermes-3-Llama-3.1-8B-AWQ

Добавлено:
solidrust/Hermes-3-Llama-3.1-8B-AWQ

— Model creator: NousResearch — Original model: Hermes-3-Llama-3.1-8B AWQ is an efficient, accurate and blazing-fast low-bit weight quantization method, currently supporting 4-bit quantization. Compared to GPTQ, it offers faster Transformers-based inference with equivalent or better quality compared to the most commonly used GPTQ settings. AWQ models are currently supported on Linux and Windows, with NVidia GPUs only. macOS users: please use GGUF models instead. — Text Generation Webui — using Loader: AutoAWQ — vLLM — version 0.2.2 or later for support for all model types. — Hugging Face Text Generation Inference (TGI) — Transformers version 4.35.0 and later, from any code or client that supports Transformers — AutoAWQ — for use from Python code

Модальности:
Генерация текста

Области применения:
Диалог / чат

Языки программирования:
Rust


Задача: Генерация текста
Автор: solidrust
Теги: llama, 4-bit, AWQ, autotrain_compatible, endpoints_compatible, conversational, text-generation-inference, awq
Лайков: 3  |  Загрузок: 35,481

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.