This is quantized version of prithivMLmods/GWQ-9B-Preview created using llama.cpp GWQ — Gemma with Questions Prev is a family of lightweight, state-of-the-art open model base from Google, built using the same research and technology employed to create the Gemini models. These models are text-to-text, decoder-only large language models, available in English, with open weights for both pre-trained and instruction-tuned variants. Gemma models are well-suited for a variety of text generation tasks, including question answering, summarization, and reasoning. GWQ is fine-tuned on the Chain of Continuous Thought Synthetic Dataset, built upon the Gemma2forCasualLM architecture. You can ensure the correct chat template is applied by using tokenizer.applychattemplate as follows: 1. Transformer-Based Design: Gemma 2 leverages the transformer architecture, utilizing self-attention mechanisms to process input text and capture contextual relationships effectively. 2. Lightweight and Efficient: It is designed to be computationally efficient, with fewer parameters compared to larger models, making it ideal for deployment on resource-constrained devices or environments. 3. Modular Layers: The…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: QuantFactory
Теги: gguf, gemma2, text-generation-inference, f16, en, endpoints_compatible, conversational
Лайков: 3 | Загрузок: 319
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.