DeepSeek-V4-Flash-0731-A100
Converted DeepSeek-V4-Flash-0731 weights for deployment on NVIDIA A100 / A800 (SM80) GPUs. > For deployment configuration, installation steps,...
Converted DeepSeek-V4-Flash-0731 weights for deployment on NVIDIA A100 / A800 (SM80) GPUs. > For deployment configuration, installation steps,...
4-bit GPTQ (W4A16) quantization of deepseek-ai/DeepSeek-V4-Flash-0731, produced with GPTQModel. This is the V1 release (first calibration run). A...
Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro,...
This model is an int4 model with group_size 128 of deepseek-ai/DeepSeek-V4-Flash-0731 generated by intel/auto-round with RTN mode. Please...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
A small ~110M parameter language model implementing the DeepSeek-V4 architecture from scratch. This is the pretrained base model...
Верная мелкомасштабная (~ 8,1B всего / ~ 2,21B активируется на токен) реплика архитектуры DeepSeek-V4, рассчитанная на обучение на...
— Base model: deepseek-ai/DeepSeek-V4-Flash — Source revision: 6e763230a9d263eca2023f1d4a5ce1bfe126cf48 — Architecture: DeepseekV4ForCausalLM — Model type: deepseekv4` — Tooling branch:...
Количество экспертов уменьшено с 256 до 192. MTP на данный момент не поддерживается. Базовый проверочный тест для версии,...
Эта 8-битная модель mlx-community/DeepSeek-V4-Pro была конвертирована в формат MLX из deepseek-ai/DeepSeek-V4-Pro. DeepSeek-V4-Pro — это модель Mixture-of-Experts с общим...