QuantFactory/Qwen2.5-14B-Instruct-1M-GGUF - Каталог нейросетей
Генерация текста

QuantFactory/Qwen2.5-14B-Instruct-1M-GGUF

Добавлено:
QuantFactory/Qwen2.5-14B-Instruct-1M-GGUF

This is quantized version of Qwen/Qwen2.5-14B-Instruct-1M created using llama.cpp Qwen2.5-1M is the long-context version of the Qwen2.5 series models, supporting a context length of up to 1M tokens. Compared to the Qwen2.5 128K version, Qwen2.5-1M demonstrates significantly improved performance in handling long-context tasks while maintaining its capability in short tasks. The model has the following features: — Type: Causal Language Models — Training Stage: Pretraining & Post-training — Architecture: transformers with RoPE, SwiGLU, RMSNorm, and Attention QKV bias — Number of Parameters: 14.7B — Number of Paramaters (Non-Embedding): 13.1B — Number of Layers: 48 — Number of Attention Heads (GQA): 40 for Q and 8 for KV — Context Length: Full 1,010,000 tokens and generation 8192 tokens — We recommend deploying with our custom vLLM, which introduce For specific guidance, refer to this section. — You can also use the previous framework that supports Qwen2.5 for inference, but accuracy degradation may occur for sequences exceeding 262,144 tokens. For more details, please refer to…

Модальности:
Генерация текста

Области применения:
Диалог / чат Следование инструкциям


Задача: Генерация текста
Автор: QuantFactory
Теги: gguf, chat, en, endpoints_compatible, conversational
Лайков: 3  |  Загрузок: 419

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.