As of 2025-09-30, create a fresh Python environment and run: For more details, please refer to vLLM documentation [[link]](https://docs.vllm.ai/projects/recipes/en/latest/DeepSeek/DeepSeek-V3_2-Exp.html) 1. Only Hopper and Blackwell data center GPUs are supported for now. 2. The kernels are mainly optimized for TP=1, so it is recommended to run this model under DP/EP mode 3. DP mode may result in increased serving latency. To mitigate this, we recommend enabling MTP to maintain optimal speed. To disable MTP, simply remove the —speculative-config flag. 4. Some users have observed improved performance on H20 machines by setting export VLLMUSEDEEPGEMM=0` We are excited to announce the official release of DeepSeek-V3.2-Exp, an experimental version of our model. As an intermediate step toward our next-generation architecture, V3.2-Exp builds upon V3.1-Terminus by introducing DeepSeek Sparse Attention—a sparse attention mechanism designed to explore and validate optimizations for training and inference efficiency in long-context scenarios. This experimental release represents our ongoing research into more efficient transformer architectures, particularly focusing on improving…
Модальности:
Генерация текста
Области применения:
Диалог / чат
Задача: Генерация текста
Автор: QuantTrio
Теги: deepseek_v32, vLLM, AWQ, conversational, endpoints_compatible, 4-bit, awq
Лайков: 4 | Загрузок: 1,552
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.