This model is quantized to 4-bit with a group size of 128. Compared to earlier quantized versions, the new quantized model demonstrates better tokens/s efficiency. This improvement comes from setting desc_act=False in the quantization configuration. We introduce the updated version of the Qwen3-4B non-thinking mode, named Qwen3-4B-Instruct-2507, featuring the following key enhancements: — Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage. — Substantial gains in long-tail knowledge coverage across multiple languages. — Markedly better alignment with user preferences in subjective and open-ended tasks, enabling more helpful responses and higher-quality text generation. — Enhanced capabilities in 256K long-context understanding. Qwen3-4B-Instruct-2507 has the following features: — Type: Causal Language Models — Training Stage: Pretraining & Post-training — Number of Parameters: 4.0B — Number of Paramaters (Non-Embedding): 3.6B — Number of Layers: 36 — Number of Attention Heads (GQA): 32 for Q and 8 for KV — Context Length: 262,144 natively. For more details, including…
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: JunHowie
Теги: qwen3, Qwen3, GPTQ, Int4, 量化修复, vLLM, conversational, text-generation-inference
Лайков: 4 | Загрузок: 2,188
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.