LlaMoE-Medium
Это модель 4x8b Llama Mixture of Experts (MoE). Обучался на OpenHermes Resort из набора данных Dolphin-2.9. Модель представляет...
Это модель 4x8b Llama Mixture of Experts (MoE). Обучался на OpenHermes Resort из набора данных Dolphin-2.9. Модель представляет...
Chat & support: TheBloke’s Discord server Want to contribute? TheBloke’s Patreon page TheBloke’s LLM work is generously supported...
Converted DeepSeek-V4-Flash-0731 weights for deployment on NVIDIA A100 / A800 (SM80) GPUs. > For deployment configuration, installation steps,...
Drop-in uncensored / abliterated weights for official DeepSeek-V4-Flash-0731 (GA), with DSpark MTP modules kept stock, for dual DGX...
A reproducible deployment and benchmark package for the official deepseek-ai/DeepSeek-V4-Flash-0731 checkpoint on two NVIDIA GB10-class nodes using vLLM...
INVALID LANGUAGE PAIR SPECIFIED. EXAMPLE: LANGPAIR=EN|IT USING 2 LETTER ISO OR RFC3066 LIKE ZH-CN. ALMOST ALL LANGUAGES SUPPORTED...
A small ~110M parameter language model implementing the DeepSeek-V4 architecture from scratch. This is the pretrained base model...
Ориентированная на выравнивание тонкая настройка DeepSeek-R1-Distill-Qwen-1.5B. Улучшение выравнивания измеряется прозрачно с использованием фреймворка Bloom от Anthropic. Три из...
DeepSeek-V3.1 — это гибридная модель, поддерживающая как режим мышления, так и режим без мышления. По сравнению с предыдущей...
DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects. — Significant improvements in benchmark performance:...