This is quantized version of Qwen/Qwen2.5-14B-Instruct-1M created using llama.cpp Qwen2.5-1M is the long-context version of the Qwen2.5 series models, supporting a context length of up to 1M tokens. Compared to the Qwen2.5 128K version, Qwen2.5-1M demonstrates significantly improved performance in handling long-context tasks while maintaining its capability in short tasks. The model has the following features: — Type: Causal Language Models — Training Stage: Pretraining & Post-training — Architecture: transformers with RoPE, SwiGLU, RMSNorm, and Attention QKV bias — Number of Parameters: 14.7B — Number of Paramaters (Non-Embedding): 13.1B — Number of Layers: 48 — Number of Attention Heads (GQA): 40 for Q and 8 for KV — Context Length: Full 1,010,000 tokens and generation 8192 tokens — We recommend deploying with our custom vLLM, which introduce For specific guidance, refer to this section. — You can also use the previous framework that supports Qwen2.5 for inference, but accuracy degradation may occur for sequences exceeding 262,144 tokens. For more details, please refer to…
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: QuantFactory
Теги: gguf, chat, en, endpoints_compatible, conversational
Лайков: 3 | Загрузок: 419
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.